Agentic Retrieval with Amazon Bedrock Knowledge Bases for Cited Insurance Claim Answers

You shouldn’t have to hunt through a dozen PDFs, emails, and scanned notes to answer “Has the estimate for claim CLM‑100482 been approved?”

That demo question teaches a simple architectural lesson: use a managed Retrieval‑Augmented Generation (RAG) pattern to ground answers in source documents, show citations for traceability, and keep humans in the loop for decisions that move money. Amazon Bedrock Knowledge Bases paired with the AgenticRetrieveStream API outline a practical approach for insurance claims workflows.

How the pieces fit

  • S3 + metadata sidecars, raw claim documents live in S3 and companion .metadata.json files expose scalar fields (policy number, date_filed, adjuster, amount) you can use to pre‑filter documents before semantic search.
  • Amazon Bedrock Knowledge Bases, a managed RAG pipeline that parses, chunks, creates embeddings, and stores vectors for you (the walkthrough used the MANAGED option for both the knowledge base and embeddings).
  • AgenticRetrieveStream, a planner, retriever, and generator loop (agentic retrieval) that can plan sub‑queries, run multiple retrieval passes, stream trace events and incremental text, and return a final answer with character‑span citations.
  • Guardrails, contextual grounding checks that compare generated output to retrieved evidence and can BLOCK outputs that miss configured grounding or relevance thresholds.

A concrete claim example to picture

The demo demonstrates a synthetic claim with this metadata for claim_id CLM-100482:

  • claim_id: CLM-100482
  • claim_type: auto
  • status: open
  • date_filed: 20260709 (the demo recommends YYYYMMDD integers)
  • amount: 14250
  • region: us-west
  • adjuster: Martha Rivera
  • policyholder: Mary Major
  • policy_number: POL-AUTO-78432
  • customer_id: CUST-MM-1042
  • household_id: HHD-MM-1042
  • document_type: adjuster_report
  • carrier: Example Insurance
  • has_subrogation: true
  • has_litigation: false
  • complexity_tier: high

The walkthrough ran in the US West (Oregon) region (us‑west‑2) and stored demo files in an example S3 bucket. The demo also used an example service role ARN. Those are demo placeholders, don’t copy example ARNs into production; create and scope your own roles and buckets.

Ingestion and metadata scoping, the mechanics that matter

Ingestion links S3 objects to a Bedrock knowledge base via a data source and an ingestion job. Each document may include a metadata sidecar (the walkthrough notes a 10 KB sidecar limit as a demo constraint, check the current Bedrock Knowledge Bases docs to confirm the limit). The demo shows using inclusion prefixes like [“claims/”] to limit what gets ingested.

The authors start ingestion with a call named start_ingestion_job on the bedrock‑agent client in their examples. SDK method names and client usage can vary across SDK versions, so confirm the exact boto3 method signature in the Bedrock docs before copy‑pasting sample code.

Why metadata sidecars are useful: pre‑filtering by metadata (equals, greaterThan, lessThanOrEquals, in/notIn, stringContains, listContains, and logical combinators like andAll/orAll) narrows candidate documents before semantic search. For example, you can filter to region = “us-west” and date_filed between 20260701 and 20260731 to scope July 2026 auto claims over $10, 000.

AgenticRetrieveStream: what to tune and expect

AgenticRetrieveStream orchestrates planning, retrieval, and generation. Important settings and behaviors the demo highlights:

  • foundationModelType, MANAGED or CUSTOM. MANAGED is convenient, CUSTOM lets you use your own foundation model configuration with more operational overhead.
  • maxAgentIteration, caps planning and retrieval rounds. The demo used 5, and the API documents a minimum allowed value of 2, so tune with that constraint in mind.
  • retriever.maxNumberOfResults, how many chunks to fetch per retrieval pass. The demo used 50 as a broad fetch (the retriever accepts values in a 1-100 range in the walkthrough examples, confirm current limits in the docs).
  • generateResponse (boolean), whether Bedrock returns a natural‑language answer for you to surface; set false when you only want traces, citations, or to run generation locally.

The API streams three event types during execution in the demo: traceEvent (planning and sub‑queries), responseEvent (incremental answer text), and result (final answer plus a results array). Citations in the final result map character spans of the answer to supporting entries in that results array; results metadata in the demo included a x-amz-bedrock-kb-source-uri field to link back to the source document. Verify field names and response JSON shapes against the current AgenticRetrieveStream reference when you implement.

Grounding, guardrails, and safety

Guardrails can enforce a grounding check that compares the LLM output to retrieved evidence and either allow or BLOCK the output when thresholds aren’t met. The demo used example thresholds such as a GROUNDING threshold of 0.85 and a RELEVANCE threshold of 0.75, with blocked messaging like “I can’t help with that request.” and “I can only answer questions using the claim records. ”

Tuning guidance: higher thresholds cut hallucination risk but increase blocked responses that require human review. Log blocked events and build a feedback loop so you can adjust thresholds, retriever breadth, or the human‑in‑the‑loop (HITL) policy.

What the demo measured, and how to read those numbers

The authors evaluated their pipeline on a small, synthetic 30‑document corpus with a 40‑question test suite. Reported outcomes (from the walkthrough) were:

  • Questions answered: 40 of 40
  • Answers carrying citations: 40 of 40
  • Mean citations per answer: 3.9
  • Expected‑source retrieval recall: 90.5%
  • Expected‑source citation recall: 81.2%
  • Chunks contradicting the requested filter: 0

On a tougher adversarial subset of 20 questions they reported retrieval recall of 96.7% and citation recall of 90.2%.

Two critical context points to keep front‑of‑mind:

  • These are demo numbers on a synthetic dataset. The authors state the evaluation used the Retrieve and RetrieveAndGenerate flows with automated grading rather than AgenticRetrieveStream, so these results are a baseline and may not predict agentic retrieval performance in the wild.
  • How the metrics were computed matters: the authors counted a “citation” when the final answer included at least one character‑span mapping to a stored chunk. Retrieval recall here means the percent of queries whose gold document appeared in the retrieved candidate set; citation recall means the percent of queries where the gold document was cited in the final answer. The walkthrough does not fully specify whether contradiction checks were automated or manual, verify the evaluation script if you plan to reproduce the benchmark.

Where teams typically trip up, and a short mitigation for each

  • Assuming synthetic performance equals production, mitigation: build a representative, anonymized test corpus from real claims and measure retrieval recall, citation precision, grounding violation rate, and latency under load.
  • Underestimating latency trade‑offs, mitigation: measure p50/p95 latency. If agentic iterations add too much delay, present an async UX (status and partial results) and tune maxAgentIteration and retriever size.
  • Poor PII/PHI hygiene, mitigation: classify and redact or tokenize PII at ingestion, encrypt with KMS, separate token vaults for mapping, and limit metadata to non‑sensitive fields unless absolutely necessary.
  • Guardrails left at defaults, mitigation: set grounding and relevance thresholds to your risk tolerance and monitor blocked events to refine thresholds and retriever recall.
  • No human‑in‑the‑loop for payments, mitigation: require human sign‑off above monetary thresholds, surface cited evidence and confidence to adjusters, and log overrides for audits.

Practical checklist before you go live (do these first)

  • Top‑3 non‑negotiables: (1) build a representative anonymized test corpus and measure recall and grounding on it; (2) require HITL sign‑off for payouts above a defined threshold; (3) classify and protect PII/PHI with redaction/tokenization, KMS encryption, and strict IAM scoping.
  • Decide MANAGED vs CUSTOM for embeddings and foundation models, MANAGED reduces ops but gives less control; CUSTOM increases control and cost and raises ops complexity.
  • Design incremental index refresh: S3 event triggers or scheduled ingestion for frequent updates; avoid full reindexes as a primary sync strategy.
  • Estimate costs for embedding storage, retrievals, and streaming inference against expected query volume and model choices.
  • Define audit policies and retention for CloudTrail logs and end‑to‑end records that link query → retrieved chunks → cited spans → human approvals.
  • Create a HITL workflow that surfaces cited source snippets, confidence, and a clear override button with mandatory audit notes.

Quick user vignette

An adjuster types: “Has the estimate for claim CLM‑100482 been approved, and when will the check be issued?” The agentic flow returns:

Yes, estimate approved 2026‑07‑12. Check issued 2026‑07‑14.

The response includes three linked evidence spans (for example): adjuster_report_CLM-100482.pdf (approval note span), payment_log.csv (issue date row), and approval_email.txt (approver signature span). Those citations let the adjuster validate the answer in seconds instead of searching multiple systems.

Key questions you should be asking, and honest, short answers

  • Can Bedrock Knowledge Bases give me cited answers for claims like CLM‑100482?

    Yes, the demo shows ingesting claim documents (S3 + metadata sidecars) into a MANAGED Bedrock knowledge base, using an agentic retrieval flow to plan and retrieve evidence, and returning answers with character‑span citations mapped to source URIs. Verify current API shapes and response fields in the Bedrock docs when implementing.

  • Will the system be correct and production‑ready out of the box?

    No, the walkthrough used synthetic data and automated grading. Real claims data contain OCR errors, inconsistent metadata, and privacy constraints. Expect to build a labeled test set, tune guardrails, and design HITL gates before rollout.

  • How do I prevent the model from inventing facts about a claim?

    Use RAG with explicit citations, enable guardrails with grounding thresholds (the demo used examples like 0.85), and BLOCK outputs below thresholds so a human reviews the case. Log blocked events and cited evidence to refine retrieval and thresholds.

  • What security controls are necessary?

    Apply least‑privilege IAM for Bedrock and S3 access, encrypt S3 and vector storage with KMS, redact or tokenize PII before ingestion, and record Bedrock calls in CloudTrail for auditing. Follow legal and compliance review for PII/PHI retention and access policies.

  • What operational unknowns must I measure before rollout?

    Measure p50/p95 latency, throughput at expected query volume, cost per query (embeddings, storage, inference), block rate under guardrails, citation precision, and index freshness under your ingestion cadence.

Final note

Agentic retrieval with Bedrock Knowledge Bases can turn scattered claim artifacts into traceable, cited answers, but the architecture is only half the job. Production safety and reliability come from real‑data evaluation, careful guardrail tuning, a clear HITL policy for money moves, robust PII handling, and operational plans for latency and cost. Treat the demo as a blueprint and validate every API, field name, and limit against the current Bedrock documentation before you copy configuration values into production.

Authors of the walkthrough: Shreya Pawaskar, Delivery Consultant, AI/ML, AWS Professional Services, and Abhishek Sharma, Senior Solutions Architect, AWS.