SageMaker agent-generated runnable benchmarks that pick the right inference instance

Pick the right SageMaker instance without guesswork: agent-generated, runnable benchmarks

Choosing instances and settings for model inference is a recurring operational drag. Teams guess, provision, and only learn later that latency missed targets or costs blew past forecasts. The aws-ai-ml skill for the Agent Toolkit for AWS addresses that bottleneck by giving MCP‑compatible coding agents the ability to produce executable SageMaker Python SDK v3 notebooks that stage models, run real load tests, compare runs, and recommend ranked deployment configurations, so engineers get repeatable, auditable guidance instead of hand-waving.

What the skill does, plainly

  • Stage models from S3, SageMaker JumpStart, or the Hugging Face Hub (including gated models with license prompts).
  • Generate runnable notebooks and scripts that execute benchmark workloads against live SageMaker endpoints, using the SageMaker Python SDK v3 plus Agent Toolkit helper calls (the announcement shows helpers such as Workload.synthetic() and start_benchmark()).
  • Compare benchmark runs and compute deltas across throughput, latency, and concurrency metrics.
  • Recommend instance types and configuration options ranked by measured performance and cost tradeoffs, and generate the deployment code you can inspect and run yourself.

Install options are flexible. You can install the aws-ai-ml skill locally through the Agent Toolkit for AWS or use a pre-configured SageMaker Studio JupyterLab image that already contains the skill and its dependencies. The Agent Toolkit manages skill discovery for MCP‑compatible coding agents such as Kiro, Claude Code, and Codex.

Safety, governance and the execution model

The skill is built to keep control in your hands. Generated code executes under your AWS credentials. Agents will ask clarifying questions and request explicit confirmation before performing potentially expensive actions such as load testing live endpoints. Installation of the skill itself does not change IAM policies in your account, but the code the agent generates requires permissions to call SageMaker and S3 APIs at runtime.

Practical governance recommendations:

  • Run benchmarks in an isolated staging account or a dedicated, tagged environment so costs are visible and contained.
  • Use a short-lived role or a role with least-privilege permissions specifically for benchmarking and staging.
  • Require human approval before the notebook actually executes the benchmark workloads.
  • Ensure any tokens (for example, Hugging Face tokens) are handled by Secrets Manager or equivalent and encrypted with KMS rather than stored in plaintext in notebooks or S3.

Quick setup checkpoints

  • Required tooling: AWS CLI 2.35+ and the Agent Toolkit for AWS. Follow the Agent Toolkit documentation for exact installation steps and prerequisites.
  • Common install commands shown in the announcement are illustrative. Follow the Agent Toolkit docs or GitHub repo to get the exact commands for your environment.
  • Example agent authentication commands in the announcement (such as a Kiro CLI login) are illustrative. Check your agent’s documentation for the authoritative CLI flags and flows.

Permissions you should expect to grant (illustrative)

The announcement notes you must run the generated code with credentials that can call SageMaker and S3 APIs. A minimal, least‑privilege benchmarking role typically needs permission to create and delete endpoint resources and to access S3 artifacts. Common actions to review and scope tightly include (verify and adapt to your policy):

  • sagemaker:CreateModel, sagemaker:CreateEndpointConfig, sagemaker:CreateEndpoint, sagemaker:DeleteEndpoint, sagemaker:DescribeEndpoint
  • sagemaker:InvokeEndpoint (if needed), sagemaker:CreateTransform (if using batch workloads)
  • s3:GetObject, s3:PutObject, s3:DeleteObject, s3:ListBucket for the staging bucket
  • iam:PassRole (if the generated code creates SageMaker resources that assume an IAM role)
  • logs:CreateLogGroup, logs:CreateLogStream, logs:PutLogEvents for logging

Make sure to scope ARNs to specific buckets and resource names where possible, and use Secrets Manager + KMS for tokens and credentials.

What the benchmarks actually measure

The aws-ai-ml-generated notebooks run real client-side load against live SageMaker endpoints and report measured metrics, not guesses. The announcement shows helper calls and code that exercise endpoints and record results. Key metrics surfaced are:

  • Throughput, requests per second and output tokens per second.
  • Latency, p50, p99, time-to-first-token (TTF), and inter-token latency.
  • Concurrency, how many simultaneous requests the setup can support under the tested workload.

Corrected example: Qwen3-1.7B vs Qwen3-8B (512/256 tokens, concurrency 4)

  • Qwen3-1.7B → Qwen3-8B: Output token throughput, 188.2 → 271.2 tokens/s (+44.1%).
  • Per-user throughput, 47.5 → 69.4 tokens/s (+45.9%).
  • Request throughput, 0.736 → 1.08 req/s (+46.7%).
  • Inter-token latency, 20.9 ms → 14 ms (−33.0%: lower is better).
  • Request latency, 5, 382 ms → 3, 658 ms (−32.0%: lower is better).
  • Time-to-first-token, 67.5 ms → 166.3 ms (+146.5%: larger model was slower to produce the first token in this run).

Hardware mapping in the example: Qwen3-8B ran on ml.g5.12xlarge (4× A10G) and Qwen3-1.7B on ml.g6.4xlarge (single L4 GPU). Keep in mind instance specs and GPU counts can vary across regions, so confirm instance details in the SageMaker/EC2 instance docs for your region.

Why did the larger model show higher token throughput but slower time-to-first-token? Possible causes include higher startup or initialization overhead, such as container startup, model loading, or inter-GPU coordination. Those factors can slow the initial token. Once streaming begins, the multi‑GPU setup and higher overall capacity can make per-token generation faster. Benchmark methodology, including warmup, batching, and concurrency patterns, also affects these measurements. Request and examine the generated notebook’s workload parameters if you care about reproducibility.

Staging models and handling gated artifacts

The skill can programmatically stage models into S3 for evaluation or deployment. For models on the Hugging Face Hub that are gated by license or terms, the agent will surface the license text and prompt you for a Hugging Face token so it can download artifacts. Before you provide tokens, confirm where they’ll be stored and whether Secrets Manager / KMS encryption is used.

JumpStart models can often be referenced directly without S3 staging, but the generated code will make clear which path it will take and what data will be written to S3.

Billing, cleanup and operational hygiene

  • Benchmarks create endpoints, run load, and store artifacts in S3, all billable. Use a tagged staging account or tags like benchmark-run-* so cost allocation is clear.
  • The announcement explicitly recommends cleanup steps: delete SageMaker endpoints created by tests, stop or delete the SageMaker Studio space used, and remove S3 objects produced by benchmark and recommendation jobs from your SageMaker AI default bucket.
  • Consider lifecycle policies on the S3 prefix used for staging and an automated endpoint teardown script to avoid accidental long-running charges.

Why this matters for engineering and finance leaders

Turning model performance decisions into reproducible, auditable artifacts shortens the path from experimentation to production. Instead of an engineer’s estimate, product managers and finance get ranked recommendations that include measured throughput and latency numbers. That visibility makes capacity planning, cost forecasting, and risk review faster and more defensible.

One caveat: benchmarks are only as useful as their workload definitions. Agents can generate and run repeatable tests for you, but someone still needs to define realistic prompts, warmup behavior, and the traffic profile representative of production.

Practical checklist before you run a benchmark

  • Use an isolated staging account or a specifically tagged environment so charges are visible and separate from production.
  • Confirm a least‑privilege IAM role with the SageMaker and S3 actions required by the generated code, and whitelist specific S3 prefixes and resource ARNs.
  • Prepare Secrets Manager entries for Hugging Face tokens and other credentials, and ensure KMS encryption is enabled.
  • Plan cleanup: record endpoint names, Studio space name, and S3 prefixes so you can delete artifacts when the run finishes.
  • Request reproducibility parameters from the agent: workload size, concurrency, warmup period, and prompt selection, and review them in the generated notebook before execution.

15-minute next steps

  1. Install the Agent Toolkit and confirm your AWS CLI (2.35+). Follow Agent Toolkit docs for the exact commands for your environment.
  2. Add the aws-ai-ml skill (via Agent Toolkit) or open the pre-configured SageMaker Studio image if available in your account.
  3. Ask the agent to generate a benchmark notebook for a non-production model and inspect the notebook for required IAM actions and S3 prefixes, do not run it yet.
  4. Create a scoped benchmarking role, prepare Secrets Manager entries, and plan a 5-minute smoke benchmark on a cheap instance to validate the workflow and cleanup steps.

Questions to ask your team or vendor

  • How are Hugging Face tokens and gated-model credentials stored and rotated?

    Prefer Secrets Manager + KMS encryption and an audit trail. Verify the agent or Studio image does not persist tokens in plaintext in notebooks or buckets.

  • Which IAM actions will the generated notebook require in our environment?

    Request a list of actions (CreateModel, CreateEndpointConfig, CreateEndpoint, InvokeEndpoint, S3 read/write/delete, iam:PassRole, etc.) and scope them to specific ARNs. Test in a least‑privilege staging role first.

  • Can we reproduce the benchmark methodology?

    Confirm warmup behavior, input sampling, tokenization settings, and concurrency ramp-up used by the generated helper calls so results are comparable across runs.

Who built this

The announcement and capability come from the Amazon SageMaker AI team, including Mona (Sr AI/ML Specialist Solutions Architect), Aryaman Kothari (Sr Product Manager), and Muzart Tuman (software engineer). The skill integrates with the Agent Toolkit for AWS and uses the SageMaker Python SDK v3 plus Agent Toolkit helper functions to produce runnable artifacts.

Key takeaways, short questions you might be asking

  • How do I add SageMaker inference expertise to my coding agent?

    Install the aws-ai-ml skill through the Agent Toolkit for AWS or use the pre-configured SageMaker Studio image; MCP-compatible agents such as Kiro, Claude Code, and Codex can then generate runnable SageMaker Python SDK v3 notebooks and scripts to stage and benchmark models.

  • Can the agent actually benchmark my endpoint?

    Yes, the agent generates notebooks that use Agent Toolkit/SageMaker helper calls (the announcement shows Workload.synthetic() and start_benchmark()) to run real load tests and report throughput, latency, and concurrency metrics. The agent prompts for explicit confirmation before running potentially expensive tests.

  • Will the agent automatically deploy production endpoints?

    No, the skill generates deployment configuration and runnable code but will not silently deploy production endpoints. The generated code runs under your AWS credentials and requires you to execute it.

  • What about gated models on Hugging Face?

    The agent can stage gated models but will surface license terms and request a Hugging Face token. Verify token handling, storage, and retention policies before providing keys in a shared environment.

  • How do I avoid surprise bills?

    Run benchmarks in a tagged staging account, inspect the generated notebook before execution, limit run durations, and delete endpoints and S3 artifacts after testing. Implement lifecycle policies and budget alerts to catch stray charges.

Agents that produce auditable, executable artifacts reduce guesswork in the model-to-production journey. The aws-ai-ml skill brings that capability into SageMaker’s domain: use it to shorten cycles and improve visibility, but pair it with least‑privilege IAM, secure token handling, reproducible workload definitions, and explicit cleanup policies.