Reads the same question against both rider documents and reports where they differ.
ClaimShield AI answers questions about insurance policy wording and shows the clause behind every answer. It is a retrieval-augmented pipeline rather than a chatbot: the model only ever sees clauses retrieved from the policy documents, and it is required to cite the ones it used.
Runs locally on a single machine: ingestion, indexing, retrieval, reranking and scoring are all local Python. The only external call is a small language model, used to phrase the answer text from the clauses it is given.
Reads the same question against both rider documents and reports where they differ.
Five stages, read top to bottom. Identifiers and numbers are as they appear in the code.
datasets/health/bajaj/health_prime/*_structured_chunks.json → load_structured_chunks() → _normalize_entry()
Structured clause JSON is the primary source; PDFs are a fallback via
load_pdfs() then chunk_documents()
(RecursiveCharacterTextSplitter, chunk_size 900 / overlap 200).
Each clause is normalised with an inferred section, section_key,
benefit_name and coverage_type (covered / not_covered / conditional / unknown).
93 clauses across 2 rider wordings — 56 individual + 37 group.
SentenceTransformer("BAAI/bge-small-en-v1.5") → faiss.IndexFlatL2(384) → indices/health/bajaj.faiss
Clauses are embedded to 384-dimension vectors and written to an exact L2 index — no IVF,
no approximation. Metadata is pickled alongside (bajaj_metadata.pkl) by
ingestion/index_builder.build_provider_index(), one index per (category, provider) pair.
classify_query_intent() + filter_chunks() → embedding_model.encode(query) → vector_store.search() → rerank()
Lightweight intent classification prunes candidates by benefit, section and limit before any
vector math. The query is embedded and searched with exact L2 over 12 pre-rank candidates
(CLAIMSHIELD_PRERANK_CANDIDATES), then the top 12 are re-ordered by the
BAAI/bge-reranker-base cross-encoder (reranker.rerank).
adaptive_top_k(min_confidence=0.45, max_chunks=8) → build_context()
Each candidate's confidence is 1 / (1 + score); clauses below 0.45 are dropped and the
result is capped at 8. build_context() serialises the survivors as
[Source i | source_file] blocks — this, and nothing else, becomes the model's input.
generate_answer() → openai/gpt-oss-20b (Groq) → [Source X] citations
Temperature 0, 1024 max tokens, reasoning_effort low. The prompt requires “answer using ONLY the provided policy context” and “cite sources using [Source X]”. If the retrieved evidence is too weak the model is never called — the endpoint returns a refusal.
The real request path through POST /search-policy, naming the functions and routes that run.
POST /search-policy (app/main.py → search_policy) validates a
SearchRequest — query, category, provider, optional product.
search.retrieve_chunks → retrieval.retrieve.retrieve_chunks calls
vector_store.load_or_create_index(category, provider). An empty index returns no results immediately.
filter_chunks(query, metadata) runs classify_query_intent and prunes candidates
by benefit, section and limit before any vector math.
embedding_model.encode(query.lower()) with bge-small-en-v1.5 (384-dim).
vector_store.search runs exact FAISS L2 over the remaining candidates, filtered by
category/product, and returns up to 12.
rerank(query, results) re-orders the top 12 with the bge-reranker-base cross-encoder.
adaptive_top_k(results, min_confidence=0.45, max_chunks=8) drops weak clauses and caps the
context at 8; the average confidence of the top 6 is computed.
build_context() assembles the context and generate_answer() calls Groq.
A spent quota surfaces as 429; a generation failure returns the evidence without an answer.
The grounding contract, as enforced in code.
[Source i | source_file] with its section, benefit, limit and waiting-period metadata,adaptive_top_k),
How a citation is produced. build_context() numbers the surviving blocks
[Source 1 … N]; the prompt requires “cite sources using [Source X]”, and every
[Source X] maps one-to-one to a retrieved clause and its file. The model cannot cite a
clause it was never given — that is the “never sees a clause it cannot cite” property, enforced by
the 0.45 floor, the 8-clause ceiling, and the one-to-one source numbering.
When retrieval finds nothing relevant. The average confidence of the top 6 falls below 0.45, the model is not called, and the answer is the refusal string “Not enough evidence found in policy documents.” with the (weak) evidence still returned for inspection.
With no API key. generate_answer() raises when GROQ_API_KEY is
unset, and the endpoint returns “Policy evidence was found, but answer generation is currently unavailable.”
with the evidence. Retrieval, reranking, adaptive filtering and the rule-based polarity classifier are all
local Python and keep working without a key — only the final phrasing needs Groq.
bajaj / healthfaiss.IndexFlatL2 exact L2 · indices/health/bajaj.faissopenai/gpt-oss-20b (Groq) · bge-reranker-base · bge-small-en-v1.5/search-policy · /compare-policies · /evaluate-retrieval · /quota · /founder-unlockapp/test_api.py)