Anthropic Sonnet 5.5: Throughput‑Optimized Claude — Measure Cost‑Per‑Task with A/B Pilots

Quick take

Anthropic released Claude Sonnet 5.5: a throughput-optimized member of the Claude 5.5 family, available on the Claude Platform and via AWS, Google Cloud/Vertex AI and Microsoft Azure. Anthropic reports big gains (faster generation, token efficiency, and higher benchmark scores), but these are vendor-reported results, and Anthropic’s system card and benchmark methodology materially affect scores and must be reviewed for independent validation.

If you’re evaluating it for production, run a targeted A/B pilot that measures cost-per-task (not just $/token), token use, latency percentiles and iteration counts before you scale.

What Sonnet 5.5 is and where you can run it

Claude Sonnet 5.5 (model ID claude-sonnet-5-5) is the second model in Anthropic’s Claude 5.5 family and is positioned as a throughput-optimized complement to Claude Opus 5.5. Anthropic makes Sonnet 5.5 available via the Claude Platform and on major cloud partners (AWS, Google Cloud / Vertex AI, Microsoft Azure). The model is closed-weights (Anthropic does not publish model weights for self-hosting).

Anthropic-confirmed specs:

  • Context window: 1, 000, 000 tokens.
  • Max interactive output: 128, 000 tokens (Anthropic documents a Message Batches beta that supports larger batch outputs; see the Sonnet 5.5 System Card for API details).
  • Reliable knowledge cutoff: June 2026.
  • Adaptive thinking: enabled by default (Anthropic documentation).
  • Effort levels: low, medium, high, xhigh, max (Anthropic lists the platform default as high. Launch materials also indicate Claude Code and some Claude apps default to medium, check the system card for surface-specific defaults).
  • Release date and lifecycle: listed as active beginning September 28, 2026. Anthropic indicates the model will remain available at least through September 28, 2027.

Pricing and billing nuances, the sticker price stayed the same

List API pricing for Sonnet 5.5 matches the Sonnet line: $2 per 1M input tokens and $10 per 1M output tokens (Anthropic pricing page). For comparison, Anthropic lists Opus 5.5 at $4 / $20 (input / output) on its pricing table.

Important billing details that change effective cost:

  • Batch API discounts: Anthropic documents a 50% Batch API discount on eligible input and output usage, so if your workload is batchable, list-rate comparisons can be misleading.
  • Cache pricing: Anthropic lists cache reads at $0.20 per 1M tokens and cache writes with tiered options (for example, a 5-minute write tier at $2.50 per 1M tokens and a 1-hour tier at $4 per 1M tokens). Which write tier you pick affects costs materially.
  • “Up to 30% lower cost per task” is Anthropic’s claim based on token and tool-call efficiency rather than a lower $/token sticker price, so your real savings depend on input/output sizes, iteration counts, tool calls, cache hit rates and whether you can use batch pricing.

Vendor-reported benchmarks and customer examples (read the system card)

Anthropic published a set of benchmark scores and customer examples in the Sonnet 5.5 launch materials and the Sonnet 5.5 System Card. These numbers are vendor-reported, and Anthropic’s benchmark prompts, effort/thinking settings, timeout rules and any tool-wrappers materially affect results. Request the system card and raw methodology files if you plan to rely on these scores for procurement decisions.

  • Terminal‑Bench 4.0: Sonnet 5.5 = 70.6% (reported at Anthropic’s default/high effort), Sonnet 5 = 10.3%, Opus 5.5 = 66.4 (Opus reported at Xhigh). (Anthropic launch materials.)
  • CursorBench 4.0: Sonnet 5.5 = 55.5%, Opus 5.5 = 57.8%, Sonnet 5 = 34.1%.
  • FrontierCode 1.1: Sonnet 5.5 = 52.1% at Xhigh, 46.2% at Max. Anthropic notes Max sometimes lowered scores because extra internal review increased runtimes and caused timeouts that FrontierCode penalizes. (For scale/context, GPT‑6 Sol scored 49.3% on the same suite in Anthropic’s charts.)
  • GDPval‑AA v2.1: Sonnet 5.5 = 1844, Opus 5.5 = 1846, Sonnet 5 = 1449. Anthropic attributes the GDPval‑AA run to Artificial Analysis in the launch materials.
  • OSWorld 2.1: Sonnet 5.5 = 80.1%, Opus 5.5 = 81.8%, Sonnet 5 = 57.0%.
  • Humanity’s Last Exam: Sonnet 5.5 = 64.5% with tools (up from 54.9% reported for Sonnet 5).

Vendor-reported customer examples (Anthropic launch materials): Balyasny Asset Management reported ~121K tokens per answer on Sonnet 5.5 vs 497K on Sonnet 5, Base44 reported 3.6 iterations per app build on Sonnet 5.5 vs 7.7 previously, and Zendesk reported processing tickets ~20% faster. These are useful signals but should be treated as vendor-reported case studies unless you obtain the underlying methodology or independent corroboration.

What Sonnet 5.5 is good for, and what you should use Opus for

Anthropic positions Sonnet 5.5 as a throughput-optimized model for well-scoped, repeatable tasks: customer support ticket handling, bug triage and fixes, polished document/slide generation, spreadsheet tasks, and other high-volume automation where speed and token efficiency matter.

Opus 5.5 remains Anthropic’s higher-capability sibling for complex, open-ended work where maximal reasoning and creative problem solving justify higher token cost and latency.

Practical evaluation guidance

  • Measure cost-per-task, not cost-per-token. Include input tokens, output tokens, iteration counts, tool calls, cache reads/writes, and batch discounts in your model.
  • Watch effort/thinking settings. Higher effort (xhigh/max) increases internal deliberation, token use and latency, which can improve quality for some tasks and worsen automated benchmark scores or time-sensitive pipelines.
  • Default effort varies by surface: Anthropic’s platform documentation lists default effort as high, while launch materials indicate Claude Code and some Claude apps default to medium. Confirm surface defaults in the system card.
  • Because the model is closed-weights, if your compliance or latency requirements demand on-prem weights you’ll need to negotiate dedicated hosting or private tenancy with Anthropic or evaluate self-hostable alternatives.

Actionable adoption checklist (run this before you commit)

  • Run a controlled A/B pilot: compare Sonnet 5.5 vs Opus 5.5 (and Sonnet 5 if relevant) on a representative slice of your workload with identical effort/thinking settings and tool wrappers.
  • Log these observability metrics per task: tokens_in, tokens_out, token_cost_per_call, tool_calls_per_task, avg_iterations_per_task, cache_hit_rate, latency_p50/p90/p99, and failure_rate.
  • Sample size & duration: run at least 2, 000 representative tasks or a minimum of two weeks (whichever captures seasonality) to surface variance and outliers.
  • Define success criteria (examples you can tune): e.g., cost-per-task ≤ 20% above current baseline, p90 latency within your SLA, and avg iterations per task ≤ 3. These are configurable to your business needs.
  • Test effort levels explicitly: evaluate Medium and High for routine workflows, and avoid Max until you validate it under your timeouts and scoring rules because extra internal review can increase runtimes and timeouts.
  • Include cache and batch scenarios in cost models: run variants with and without caching, and model batch discounts for high-volume pipelines.
  • Ask the vendor for the Sonnet 5.5 System Card and full benchmark methodology (prompts, seeds, timeout rules, effort settings, runtime environment, and tool wrappers) before using vendor benchmark numbers in procurement decisions.
  • If security/compliance is required, request details about Anthropic’s stated “cyber safeguards and reasoning-extraction classifiers” and ask for any red-team or third-party audit reports.

Architectural note, hybrid patterns still win on cost

For most production pipelines the sensible pattern is hybrid: use smaller, cheap decision/routing models to classify or route inputs, then call Sonnet 5.5 or Opus 5.5 only for generation or complex reasoning. Decision models reduce calls to the generator and lower cost for high-volume operations, and Sonnet 5.5 handles the heavier generation work when necessary.

Caveats and the verification gap

Two points every buyer should treat as gating. First, the benchmark numbers and customer ROI examples in Anthropic’s materials are vendor-reported. Benchmarking outcomes are sensitive to prompts, effort/thinking settings, timeout rules and tool wrappers, so request the system card and raw methodology to reproduce results. Second, Anthropic locks certain sampling parameters (raising a 400 error if you change temperature/top_p/top_k), which improves reproducibility but limits creative sampling adjustments.

On the weird bits you should ask about

  • The “first Sonnet to beat Pokémon Red using screenshots” line appeared in launch materials, so ask Anthropic for the dataset, rules and full run logs if that example matters to your evaluation.
  • Anthropic reports “>30% faster” and “up to 30% lower cost per task” versus Sonnet 5, so request absolute latency metrics (ms/token or tokens/sec), the hardware used for measurements, and the cost-per-task worked examples behind those percentages.
  • Confirm which cache write tier (5m vs 1h) was used in any vendor cost examples and whether batch discounts were applied to the customer case studies you care about.

Quick worked example, how token pricing converts to task cost

This is an illustrative calculation using Anthropic’s list rates (no batch or cache effects): assume 2, 000 input tokens and 1, 000 output tokens in one call.

  • Input cost: 2, 000 / 1, 000, 000 * $2 = $0.004
  • Output cost: 1, 000 / 1, 000, 000 * $10 = $0.010
  • Total per call (list rates): ~$0.014

If Sonnet 5.5 reduces iterations or output size versus another model, the per-task cost drops proportionally. Include cache reads/writes and batch discounts in your own modeling for realistic estimates.

Key questions readers ask, short answers

  • Is Sonnet 5.5 available to run today and where?

    Yes. Anthropic lists Sonnet 5.5 (model ID claude-sonnet-5-5) as available on the Claude Platform and via cloud partners (AWS, Google Cloud / Vertex AI, Microsoft Azure). It is a closed-weights model and must be used through Anthropic or its cloud integrations.

  • Did Anthropic change the $/token price?

    No. List pricing remained $2 per 1M input tokens and $10 per 1M output tokens. Anthropic points to token and tool-call efficiency (plus batch discounts) as the path to lower cost per task.

  • Are the benchmark gains independently verified?

    The published scores and customer examples are vendor-reported in Anthropic’s launch materials and system card. Request the system card and raw methodology or run your own replication to validate claims for your workloads.

  • Which model should I use in my pipeline: Sonnet or Opus?

    Use Sonnet 5.5 for well-scoped, high-throughput tasks that value latency and token efficiency. Use Opus 5.5 for complex, open-ended problems where higher capability is needed. Many teams combine a decision/routing model with Sonnet/Opus in a hybrid architecture.

  • Do I need to change how my agents and tools call the model?

    Yes, pay attention to adaptive thinking and effort defaults. Anthropic defaults differ by surface (platform default = high; some Code/apps default to medium). Also, certain sampling parameters are locked and attempting to change them returns an API error; follow Anthropic’s migration guide for correct thinking syntax and effort settings.

Final note

Sonnet 5.5 is Anthropic’s throughput play: large context, faster generation and vendor-reported token efficiency that promise real savings for high-volume, well-scoped automation. Those vendor-reported gains are worth testing, not accepting on faith, so pull the Sonnet 5.5 System Card, replicate the runs that matter to you, and run an A/B pilot that captures token efficiency, latency percentiles and iteration counts before you scale.