An agent recommends 1, 500 units, do you trust the number or demand the math?
An agent answers with conviction: “Recommending 1, 500 units because demand forecast is 1, 200 and we target a 1.25× safety factor.” Nice sentence. But if you run the business, the real question is whether the agent actually used the forecast, respected budget and warehouse limits, and documented the trade‑offs that produced 1, 500.
If you can’t verify tool selection, workflow steps, and the arithmetic behind a recommendation, that fluent answer is a human‑level operational risk. The sample reference implementation built on Amazon Bedrock AgentCore and the Strands Agents SDK shows how to measure not just helpfulness, but the full decision pipeline: tool calls, constraint checks, KPI impact, and explainability.
Three quick moves for leaders
- Clone the supply‑chain sample and run it in a sandbox to see traces and evaluations end‑to‑end (links below).
- Create one Constraint Satisfaction evaluator and gate it in CI for every agent change (start strict, tune later).
- Enable 0.5% online sampling with CloudWatch alarms and a human review queue for initial calibration.
A practical evaluation framework, three stacked lenses
Evaluate agentic systems across three layers so you don’t mistake polish for correctness:
- Layer 1, Built‑in evaluators (fast, generic): standard checks such as Helpfulness, Tool Selection Accuracy, Response Relevance, Instruction Following, Faithfulness, Builtin.Correctness, and Builtin.GoalSuccessRate.
- Layer 2, Custom business evaluators (domain rules): constraint checks, KPI attainment, SQL correctness, route feasibility, recommendation groundedness, etc.
- Layer 3, Explainability evaluators (cross‑cutting): Decision Rationale Quality, Evidence Attribution, Constraint Reasoning, Trade‑off Explanation, Tool‑use Explainability, and Assumption Disclosure.
Layering keeps fast, repeatable checks separate from domain logic and explainability. Teams can run automated gates while preserving human‑readable evidence for audits.
Reference scenario: AnyCompany Retail and the multi‑agent topology
The sample models a fictitious multinational retailer, AnyCompany Retail, and uses a modular topology: one orchestrator agent delegates to four specialized sub‑agents, optimization, distribution, routing, and analytics. That split mirrors real workflows where different skills and tools must cooperate.
Implementation highlights (as implemented in the sample): Strands Agents SDK for coordination; Amazon Bedrock AgentCore runtime, memory, and Observability for execution and traces; MCP tools served via an AgentCore MCP Server backed by mock endpoints on Amazon API Gateway; Amazon RDS as a grounding target in evaluator examples; and CloudWatch for streaming traces and alerts.
Why the topology matters
Correctness is a workflow property, not just a text property. A plausible recommendation can still be wrong if the orchestrator routed to the wrong tool, a sub‑agent used stale RDS data, or a constraint check was skipped. Observability traces plus layered evaluators let you reconstruct the decision path and prove whether the math was actually performed.
Run the sample and its evaluators (concise)
Download the supply‑chain example: https://github.com/aws-samples/sample-agentic-genai-agentcore/tree/main/aws-genai-evaluations-supply-chain. Additional samples are at https://github.com/awslabs/agentcore-samples/tree/main/06-workshops/07-AgentCore-evaluations.
Deploy with Terraform (follow the repo README for details):
cd terraform
cp terraform.tfvars.example terraform.tfvars
terraform init
terraform apply
Terraform outputs include runtime ARNs, the AnyCompany Retail API URL, the evaluator API URL, and the memory ARN. Clean up with terraform destroy.
Exercise the agent using the provided test client (lightweight functional tests):
python test_agent.py –runtime-arn “<supply_chain_arn>” –region <your_region>
python test_agent.py –runtime-arn “<supply_chain_arn>” –category optimization –region <your_region>
Register and run evaluators with the test helper:
python test_evaluator.py –api-url “https://<evaluators-api-url>” create
python test_evaluator.py –api-url “https://<evaluators-api-url>” run –agent-id “supply_chain_orchestrator_agent-<id>” –session-id “<session-id>” –evaluators “sc_optimization_constraint-<id>, sc_distribution_groundedness-<id>, Builtin.Correctness, Builtin.GoalSuccessRate”
The API returns HTTP 202 immediately and saves results asynchronously to S3 at a path formatted as:
s3://<amzn-s3-demo-agent-source-bucket>/evaluations/<session-id>/<timestamp>/EvaluationResults.md
To remove evaluators:
python test_evaluator.py –api-url “https://<evaluators-api-url>” delete –evaluator-ids “sc_optimization_constraint-<id>, sc_distribution_groundedness-<id>”
Prerequisites (exact versions required by the sample): AWS CLI; AWS SAM CLI v1.100.0+; Docker v20.x+; Node.js v18.x+; Python v3.11+. Dockerfile dependencies include strands-agents, strands-agents-tools, requests, bedrock-agentcore, and boto3.
On‑demand vs online evaluation, when to use each
- On‑demand: replay recorded sessions for development, benchmarking, and CI gates. Deterministic and ideal for regression tests.
- Online: continuous, sampled monitoring in production. OnlineEvaluationConfig can reference evaluator ARNs and specify a sampling rate (the sample suggests 1-10% as a starting range).
In online mode evaluators can read traces from AgentCore Observability and stream results to CloudWatch for dashboards and alarms. Online evaluation increases trace volume and evaluator compute. Pilot sampling and measure cost and latency before scaling.
Concrete evaluators you can copy into a pipeline
Built‑in evaluators included in the sample (names shown in examples): Helpfulness; Tool Selection Accuracy; Response Relevance; Instruction Following; Faithfulness; Builtin.Correctness; Builtin.GoalSuccessRate.
Custom evaluators in the supply‑chain demo include:
- Plan Coherence
- Tool Trajectory
- Constraint Satisfaction (Budget; Inventory coverage; Warehouse capacity)
- KPI Attainment (fill‑rate, revenue improvement)
- Recommendation Groundedness
- Risk Impact
- Route Feasibility and SLA checks
- SQL Correctness; Data‑grounding; No Unsupported Claims
Constraint checks in the sample are explicit and intentional so you can automate them. For example:
- Inventory coverage: demand ≤ qty ≤ demand × safety_factor (the sample uses safety_factor = 2 by default; make this configurable)
- Budget check: assert (budget_limit − budget_used) ≥ proposed_cost
- Warehouse capacity: recommended_qty ≤ capacity
Those rules are a starting point, surface them as configuration so product owners and planners can tune safety_factor and budget policies per SKU or business unit.
Explainability as a first‑class metric
Explainability evaluators verify that an agent didn’t just state a number, but exposed the inputs and reasoning. The sample uses exemplars such as:
“Recommending 1, 500 units because demand forecast is 1, 200 and we target a 1.25× safety factor”
“Budget allows up to 1, 800 units but warehouse capacity limits us to 1, 600, so we recommend 1, 500”
Free‑text rationales are useful for humans but hard to verify automatically. The recommendation: instrument agents to emit structured rationales (a small, machine‑readable payload that evaluators can assert against). For example, an explainability payload might include fields like this:
DecisionRationale: { “forecast”: 1200, “safety_factor”: 1.25, “computed_qty”: 1500, “budget_limit”: 1800, “warehouse_capacity”: 1600, “constraint_violations”: [] }
When agents emit structured rationales, explainability evaluators can perform deterministic checks (math, referenced data IDs, which tool was used) instead of brittle natural language matching.
Operational safety: Guardrails and Observability
Evaluations measure behavior. Guardrails enforce it at runtime. Amazon Bedrock Guardrails (as implemented in the sample) provide runtime safeguards, content filters, denied‑topic detection, grounding validation, while AgentCore Evaluations scores adherence and surfaces regressions.
AgentCore Observability collects traces of tool calls, SQL queries, and decision paths. The sample routes traces to Amazon CloudWatch so evaluators, dashboards, and alarms tie back to concrete sessions and S3 evaluation artifacts.
How this fits into engineering workflows
Two practical patterns to adopt quickly:
- CI gating: run on‑demand evaluators in pull request or build pipelines. Example rule, fail builds if Constraint Satisfaction fails for more than 1% of regression sessions or if Builtin.Correctness score drops below your CI threshold.
- Production sampling: enable online evaluation sampling (pilot 0.5%, 1%). Route evaluation failures to CloudWatch alarms and a human review queue for the first 1, 000 sampled failures so your team can tune evaluator rules and reduce false positives.
Suggested starter thresholds (teams must adapt these to domain risk):
- Tool Selection Accuracy ≥ 95% for admin releases.
- Builtin.Correctness ≥ 99% for critical decision paths in CI.
- Constraint Satisfaction = 100% for rules that protect safety or compliance; non‑blocking warnings for softer business rules.
- Alert if sampled sessions show >5% ungrounded recommendations in a rolling 24‑hour window (trigger P1 review).
These are conservative starting points. The sample intentionally leaves numeric cutoffs to customers. Use the pilot period to collect real data and iterate thresholds rather than trusting any single number off the shelf.
Practical limits, known gaps, and how to address them
- No out‑of‑the‑box thresholds. The reference defines checks but not pass/fail cutoffs. Recommendation: start strict in CI, collect production samples, then loosen thresholds where acceptable.
- Cost and latency of online evaluations. The sample suggests 1-10% sampling. Recommendation: run a cost experiment with 0.1%, 0.5%, and 1% for two weeks to measure trace storage and evaluator compute before scaling.
- False positives/negatives in evaluators. Recommendation: queue initial failures for human review (first 1, 000 cases) and use that feedback to refine rules and examples.
- Mock endpoints in the sample. The demo uses API Gateway mocks. In production replace mocks with secure internal APIs, VPC endpoints, least‑privilege IAM, and PII‑masking before traces leave your network.
- Human‑in‑the‑loop calibration. Recommendation: capture reviewer labels and feed them into evaluator rule updates or model fine‑tuning pipelines.
What to try next (a compact starter playbook)
- Clone the repo and run the sample in an isolated AWS sandbox. Inspect traces in CloudWatch and results in S3.
- Author a Constraint Satisfaction evaluator for a single SKU (demand, safety_factor, budget, capacity) and add it to CI gates.
- Enable 0.5% online sampling, route failures to CloudWatch, and have a human review queue for the first 1, 000 failures to calibrate evaluators.
Key questions a curious leader will ask, and honest answers
-
Can I evaluate an agent’s decisions, not just its text?
Yes. AgentCore Evaluations supports built‑in evaluators for response quality and lets you register custom evaluators that check business constraints, SQL correctness, groundedness, and KPIs. Explainability evaluators verify the agent’s rationale independently. Start by requiring structured rationales from agents so evaluators can perform deterministic checks.
-
How do I run these evaluations in development and production?
Use on‑demand mode to replay sessions in CI and for regression testing. Use online mode for continuous monitoring with sampling; OnlineEvaluationConfig accepts evaluator ARNs and a sampling rate (the sample suggests 1-10%, but pilot smaller first).
-
Where do evaluation results go and how are they audited?
Evaluation outputs are written asynchronously to S3 at s3://<amzn-s3-demo-agent-source-bucket>/evaluations/<session-id>/<timestamp>/EvaluationResults.md, and traces stream to Amazon CloudWatch. Combine these artifacts for an auditable trail; implement retry/dead‑letter handling for async failures.
-
Do I still need Guardrails if I run Evaluations?
Yes. Guardrails enforce runtime constraints (content filters, grounding checks) while Evaluations measures behavior and surfaces regressions. They are complementary controls: one enforces safety, the other proves compliance and detects drift.
-
Where can I get the sample implementation to try this end‑to‑end?
Download the supply‑chain evaluation sample at https://github.com/aws-samples/sample-agentic-genai-agentcore/tree/main/aws-genai-evaluations-supply-chain. Additional examples are at https://github.com/awslabs/agentcore-samples/tree/main/06-workshops/07-AgentCore-evaluations.
Final note
Evaluating multi‑agent systems means treating decisions as auditable workflows: record tool calls, require structured rationales where possible, and score both business correctness and explainability. The reference implementation and Evaluations workflow demonstrated by Kanishk Mahajan (Principal, AI/ML with AWS Professional Services) provides a hands‑on starting point, Terraform deployment, test clients, evaluator scripts, and examples you can adapt for CI/CD and production monitoring. Use the sample to build a small, measurable feedback loop before expanding to broader fleets of agents.