CS ClaimShield AI live demo

ClaimShield AI answers questions about insurance policy wording and shows the clause behind every answer. It is a retrieval-augmented pipeline rather than a chatbot: the model only ever sees clauses retrieved from the policy documents, and it is required to cite the ones it used.

Runs locally on a single machine: ingestion, indexing, retrieval, reranking and scoring are all local Python. The only external call is a small language model, used to phrase the answer text from the clauses it is given.

Try:

Retrieval pipeline

Five stages, read top to bottom. Identifiers and numbers are as they appear in the code.

  1. 01 ingest

    datasets/health/bajaj/health_prime/*_structured_chunks.json → load_structured_chunks() → _normalize_entry()

    Structured clause JSON is the primary source; PDFs are a fallback via load_pdfs() then chunk_documents() (RecursiveCharacterTextSplitter, chunk_size 900 / overlap 200). Each clause is normalised with an inferred section, section_key, benefit_name and coverage_type (covered / not_covered / conditional / unknown). 93 clauses across 2 rider wordings — 56 individual + 37 group.

  2. 02 index

    SentenceTransformer("BAAI/bge-small-en-v1.5") → faiss.IndexFlatL2(384) → indices/health/bajaj.faiss

    Clauses are embedded to 384-dimension vectors and written to an exact L2 index — no IVF, no approximation. Metadata is pickled alongside (bajaj_metadata.pkl) by ingestion/index_builder.build_provider_index(), one index per (category, provider) pair.

  3. 03 retrieve

    classify_query_intent() + filter_chunks() → embedding_model.encode(query) → vector_store.search() → rerank()

    Lightweight intent classification prunes candidates by benefit, section and limit before any vector math. The query is embedded and searched with exact L2 over 12 pre-rank candidates (CLAIMSHIELD_PRERANK_CANDIDATES), then the top 12 are re-ordered by the BAAI/bge-reranker-base cross-encoder (reranker.rerank).

  4. 04 ground

    adaptive_top_k(min_confidence=0.45, max_chunks=8) → build_context()

    Each candidate's confidence is 1 / (1 + score); clauses below 0.45 are dropped and the result is capped at 8. build_context() serialises the survivors as [Source i | source_file] blocks — this, and nothing else, becomes the model's input.

  5. 05 answer

    generate_answer() → openai/gpt-oss-20b (Groq) → [Source X] citations

    Temperature 0, 1024 max tokens, reasoning_effort low. The prompt requires “answer using ONLY the provided policy context” and “cite sources using [Source X]”. If the retrieved evidence is too weak the model is never called — the endpoint returns a refusal.

What happens on a question

The real request path through POST /search-policy, naming the functions and routes that run.

  1. Route. POST /search-policy (app/main.py → search_policy) validates a SearchRequest — query, category, provider, optional product.
  2. Load the index. search.retrieve_chunks → retrieval.retrieve.retrieve_chunks calls vector_store.load_or_create_index(category, provider). An empty index returns no results immediately.
  3. Intent filter. filter_chunks(query, metadata) runs classify_query_intent and prunes candidates by benefit, section and limit before any vector math.
  4. Embed the query. embedding_model.encode(query.lower()) with bge-small-en-v1.5 (384-dim).
  5. Vector search. vector_store.search runs exact FAISS L2 over the remaining candidates, filtered by category/product, and returns up to 12.
  6. Rerank. rerank(query, results) re-orders the top 12 with the bge-reranker-base cross-encoder.
  7. Threshold & assemble. adaptive_top_k(results, min_confidence=0.45, max_chunks=8) drops weak clauses and caps the context at 8; the average confidence of the top 6 is computed.
  8. Answer or refuse. If average confidence is below 0.45 the endpoint returns an evidence-only refusal — “Not enough evidence found in policy documents.” — without calling the model. Otherwise build_context() assembles the context and generate_answer() calls Groq. A spent quota surfaces as 429; a generation failure returns the evidence without an answer.

Why the answers are trustworthy

The grounding contract, as enforced in code.

The model is given

  • at most 8 retrieved clauses that passed the 0.45 confidence floor,
  • each labelled [Source i | source_file] with its section, benefit, limit and waiting-period metadata,
  • nothing else.

The model is not given

  • clauses that scored below 0.45 (dropped by adaptive_top_k),
  • any document that retrieval did not return,
  • the full policy text, conversation history, or any outside knowledge.

How a citation is produced. build_context() numbers the surviving blocks [Source 1 … N]; the prompt requires “cite sources using [Source X]”, and every [Source X] maps one-to-one to a retrieved clause and its file. The model cannot cite a clause it was never given — that is the “never sees a clause it cannot cite” property, enforced by the 0.45 floor, the 8-clause ceiling, and the one-to-one source numbering.

When retrieval finds nothing relevant. The average confidence of the top 6 falls below 0.45, the model is not called, and the answer is the refusal string “Not enough evidence found in policy documents.” with the (weak) evidence still returned for inspection.

With no API key. generate_answer() raises when GROQ_API_KEY is unset, and the endpoint returns “Policy evidence was found, but answer generation is currently unavailable.” with the evidence. Retrieval, reranking, adaptive filtering and the rule-based polarity classifier are all local Python and keep working without a key — only the final phrasing needs Groq.

Stack

Corpus
2 rider wordings · 93 clauses (56 individual + 37 group) · bajaj / health
Index
93 × 384-dim · faiss.IndexFlatL2 exact L2 · indices/health/bajaj.faiss
Models
openai/gpt-oss-20b (Groq) · bge-reranker-base · bge-small-en-v1.5
Endpoints
/search-policy · /compare-policies · /evaluate-retrieval · /quota · /founder-unlock
Evaluation
29 hand-labelled queries — section match, semantic answer similarity, MRR, polarity confusion matrix
Code
26 modules · 3,312 lines of Python · 2 API tests (app/test_api.py)