When a four‑point tweak reshuffles the summit: what Artificial Analysis v4.2 means for buyers
Artificial Analysis released Intelligence Index v4.2 and raised GPT‑6 Astra by four points. The leaderboard changed: Anthropic’s Claude Fable 5.1 remains first, Astra moved to second, and Meta sits third. That math is small. The signal is not. The update changes what the index rewards, how resistant it is to gaming, and how easy it will be for you to validate vendor claims, all things that matter when you’re buying AI for business automation or deploying continuous agents.
What changed in v4.2, the essentials
- Added two benchmarks: AA‑Briefcase (real‑world knowledge‑work tasks) and GDP.pdf (Surge AI’s PDF document analysis test), according to Artificial Analysis’ v4.2 materials.
- Removed GPQA‑Diamond from the active evaluation list (GPQA‑Diamond does not appear among the ten evaluations AA lists for v4.2).
- Increased private test weighting to 40% of the index signal, as stated in AA’s v4.2 methodology notes.
- Reported fixes for prior scoring errors and grading tweaks intended to stabilize results.
- Leaderboard after v4.2: Claude Fable 5.1 (Anthropic) is 1st, GPT‑6 Astra is 2nd (a four‑point gain), and Meta is 3rd, per AA’s published leaderboard.
- Artificial Analysis says Version 5 has been in development for roughly eight months and will roll out in stages, as reported alongside the v4.2 announcement and summarized by Matthias Bastian (The Decoder, Sep 5, 2026).
Vendor checklist, the questions procurement should ask first
Before you get lost in head‑to‑head rankings, ask vendors for numbers you can run against your workflows. A short checklist you can paste into an RFP or a vendor call:
- Provide tokens‑per‑task for three representative workflows (short, medium, long) and the raw prompt/response pairs (redacted for sensitive data).
- Report average model calls (actions) per workflow and the harness or evaluation wrapper used to measure them.
- Share the harness recipe or open‑source harness used for testing, or allow a trusted third party to run the same harness against a redacted sample.
- Disclose cost assumptions for cost‑to‑performance comparisons (token prices, cache hit rates, tool invocation costs, and infrastructure assumptions).
- Provide a redacted sample of any private tests or permit an independent audit under NDA. If unavailable, explain why and provide an auditable substitute.
Why those questions matter: harnesses, private tests, and action efficiency
Benchmarks differ because they test different things. Three factors explain why one index crowns a leader while another ranks the same model lower:
- Harness (the test wrapper): the harness formats tasks, controls state and caching, and decides how the model can reuse internal reasoning. ARC‑AGI shows harness choice can flip results dramatically. They reported GPT‑6 Astra scoring 62.7% on ARC‑AGI‑3 Semi‑Private with the Standard harness (cost ≈ $26, 098) versus scores in the high‑90s (98.6%, 99.9%) with the Provider Adapter harness at roughly $17k, $19k, depending on configuration. Those are ARC‑AGI’s published numbers and they illustrate how evaluation wrappers change perceived capability and cost.
- Private vs public tests: AA’s move to make private tests 40% of the index is meant to reduce gaming of public prompt banks. Private tests catch overfitting, but they also reduce outside reproducibility. Buyers need compensating transparency, like redacted samples or third‑party audits, to trust private results.
- Action efficiency: how many model calls (actions) a workflow needs. ARC‑AGI reports Astra used fewer actions than the human baseline on 96.0% of levels and 51.7% fewer actions per level on average, metrics ARC uses to quantify action efficiency. Fewer actions mean fewer model calls, fewer tokens, lower per‑task cost, and lower latency. Artificial Analysis similarly reports Astra used fewer tokens per task than other tested frontier models in their sample.
What AA’s methodological changes reward
By adding AA‑Briefcase and GDP.pdf, AA shifts the index toward multi‑step, document‑heavy, professional workflows, the kinds of tasks buyers actually pay for: contract review, analyst synthesis, sales sequences, and other knowledge‑work automation. Removing GPQA‑Diamond drops a checklist item that no longer differentiates frontier models. Raising private test weighting penalizes models tuned only to public prompt banks and rewards robustness on unseen, realistic tasks.
That is a deliberate editorial judgment by AA: make the index harder to game and more reflective of autonomous, document‑centric work. The trade‑off is transparency. If 40% of a score comes from nonpublic tasks, external auditors and customers need compensating controls to maintain trust.
Operational metrics that matter for AI automation
Leaderboards shape market narratives. Operational metrics tell you what you will actually pay for and how the model will behave in production. Prioritize these:
- Actions per task (model calls): Ask for average and tail counts. For continuous agents, a reduction in calls is a direct cost and latency win.
- Tokens per task: Confirm how tokens were counted, including input, reasoning, cache writes, and answers. AA reports Astra used fewer tokens per task in its v4.2 sample. Get the raw table.
- Cost normalization: Make vendors state the token pricing and caching assumptions they used for any cost‑to‑performance claims. Artificial Analysis provides a cost model, but its inputs matter.
- Task fidelity for your documents: If your workflows are PDF or contract heavy, benchmarks like GDP.pdf matter. Insist on real‑world document tests or a redacted sample you can validate.
- Reproducibility: If private tests dominate the score, require redacted samples, an independent audit, or the ability to run the same harness against your redacted data.
How to interpret the v4.2 leaderboard, and where to be skeptical
Astra’s reported four‑point gain moved the needle enough to shift narratives. Still, watch for a few caution flags:
- Absolute significance depends on the scale. A four‑point jump matters more on a compressed scale than on a very large one. AA does not publish every before/after delta in a single public table. Ask AA for the leaderboard CSV or changelog to see absolute scores and deltas per model.
- “Astra used fewer tokens per task than other frontier models” is AA’s reported finding on their sampled tasks. Get the per‑task breakdown rather than relying on the summary statement.
- AA’s phrasing that Anthropic, OpenAI, Meta, and Zhipu AI “all share the lead” on cost‑to‑performance reflects AA’s cost model and charting. Different cost assumptions would change the Pareto front. Treat cost claims as AA‑model specific unless vendors disclose the pricing inputs.
- Other evaluators have produced different leaderboards. Epoch AI and ARC‑AGI reported different orderings and point scales. Those rankings use different benchmark mixes and harnesses, so point totals are not directly comparable without normalization.
Practical procurement playbook
When a vendor hands you a leaderboard screenshot, act quickly and decisively:
- Ask for the numbers behind the chart, including the full leaderboard CSV, the weighting breakdown (public vs private), and the exact tests added or removed.
- Run a short acceptance test: three representative workflows, measured tokens, actions, and latency using your prompts and a redacted data sample.
- Require cost transparency. Vendors should provide cost‑to‑serve calculations using your token pricing and expected cache characteristics.
- Insist on auditability. If an index relies heavily on private tests, require a redacted sample or an independent audit before you make a long‑term contract decision based on that index.
Key questions, short answers and actions
-
Did Artificial Analysis change its index because Astra’s score was questioned?
Yes. Artificial Analysis released v4.2 after criticism that its earlier ranking underplayed Astra’s gains; AA’s v4.2 materials and reporting by Matthias Bastian (The Decoder, Sep 5, 2026) document the methodological changes and Astra’s four‑point rise. Action: request AA’s changelog and the before/after leaderboard CSV.
-
Does Astra now lead the pack?
No. In AA’s v4.2 leaderboard Claude Fable 5.1 (Anthropic) remains first, GPT‑6 Astra is second, and Meta is third. Other evaluators (Epoch AI, ARC‑AGI) have reported different orderings under different methodologies and harnesses; compare apples to apples before deciding.
-
Are the index changes substantive or cosmetic?
Substantive: AA added real‑world and PDF benchmarks, removed a less‑differentiating test, increased private‑test weighting to 40%, and reported scoring fixes and grading tweaks, each can meaningfully reshape rankings and which capabilities are rewarded. Action: ask AA which fixes affected which benchmarks and by how much.
-
Do private tests make rankings more reliable?
They reduce gaming but decrease transparency. Private tests catch overfitting to public prompts, but buyers should demand redacted samples or an independent audit before relying on private‑heavy rankings for procurement.
-
Which operational metrics should drive buying decisions?
Action efficiency (model calls per task), tokens‑per‑task, and cost‑to‑performance under your pricing assumptions. ARC‑AGI’s published harness comparisons show
Further reading
If you want a concise primer on how language‑model evaluations are structured and compared, this overview is a useful starting point.