Gemini 4 Argon, Google narrows the frontier gap but don’t swap vendors yet
In short: Argon meaningfully closes capability gaps with top OpenAI and Anthropic models, and its 1, 000, 000-token context/output limits plus Long-Decode Continuation make new long-form workflows possible. But Argon’s higher average output-token use and several unresolved runtime and billing details mean most enterprises should pilot before replacing existing providers.
Executive signal for leaders
What to know immediately: Google/DeepMind’s Gemini 4 Argon brings unprecedented scale (1, 000, 000 tokens for input and output) and promising gains on agentic tasks and human preference tests. Independent evaluators report cost advantages under Google’s promotional pricing, but Argon tends to generate many more output tokens per task, which can erase that edge once promo rates end. Prioritize sandbox testing, track token usage, and insist on clear billing and continuation rules before moving production workloads.
Key capabilities, concise and practical
- Context and output scale: Google reports a 1, 000, 000-token input context window and up to 1, 000, 000 output tokens, large enough to hold extensive document collections, long transcripts, or sustained multi-step reasoning in a single session.
- Multimodal inputs: Google says Argon accepts text, images, video, and audio as inputs; Google also states outputs are text (vendor claim).
- Long Decode Continuation: a new API feature meant to pause long responses and resume them via follow-up calls, intended to avoid timeouts during extended generation. Implementation details, such as state handling, billing per resumed segment, and latency tradeoffs, have not been fully disclosed by Google.
- Phased rollout and early access: Argon is initially available to “trusted cyber defenders” through the Fairwind program and to certain U.S. government pilots. Google describes a “phased approach” and says it will expand access to paying API customers and Google AI Ultra subscribers “as soon as possible.”
“trusted cyber defenders”
Pricing, standardized units and the cache math
Google published promotional and regular list rates that are aggressive for a frontier model. Presented here as $/million tokens:
- Gemini 4 Argon (promotional): $2 per million input tokens; $10 per million output tokens; cache read operations ≈ $0.10 per million tokens (calculated as 95% cheaper than the $2 input token price); cache write operations: not disclosed.
- Gemini 4 Argon (regular): $4 per million input tokens; $20 per million output tokens; cache read operations ≈ $0.20 per million tokens (95% discount on $4); cache write operations: not disclosed.
- GPT‑6 Astra: $10 per million input tokens; $50 per million output tokens; cache read operations $1 per million; cache write operations $12.50 per million.
- Claude Fable 5.1: $10 per million input tokens; $50 per million output tokens; cache read operations $0.25 per million; cache write operations $12.50 (5 min.) / $20 (1 hr.).
- Claude Opus 5.5: $4 per million input tokens; $20 per million output tokens; cache read operations $0.20 per million; cache write operations $5 (5 min.) / $8 (1 hr.).
Footnote: Google does not explicitly list the cache read figure but states the cache is 95% cheaper than the regular input price. The $0.10 and $0.20 numbers above follow directly from that 95% reduction.
Benchmarks and real numbers, what the evaluators report
Independent and vendor evaluations paint a mixed, mostly encouraging picture. Below are notable scores and their sources. Methodologies vary, and several evaluators’ prompt and generation settings were not publicly available for full replication.
- Artificial Analysis, Intelligence Index: Argon (High) scored 53 points and ties with GPT‑6 Astra (max); Argon improved +23 points over Gemini 3.1 Pro Preview (Artificial Analysis results cited by independent reporting). Artificial Analysis also reports Argon averaged about 62, 000 output tokens per Intelligence Index task versus GPT‑6 Astra’s 27, 000 output tokens, a divergence that materially affects per-task billing under tokenized pricing.
- AutomationBench‑AA / Terminal Bench 4: Argon scored 77.5% on AutomationBench‑AA (first place) and 57% on Terminal Bench 4 (big jump over Gemini 3.1 but behind some Anthropic/OpenAI models). These are the Artificial Analysis agentic and task automation variants.
- AA‑Omniscience (hallucination & accuracy): Artificial Analysis reports Argon’s hallucination rate at 15% on AA‑Omniscience, and an accuracy of 50% (versus GPT‑6 Astra’s 63% accuracy and higher hallucination rates reported for some GPT‑6 variants). Definitions of “hallucination” and “accuracy” differ by benchmark; consult the evaluator’s methodology for specifics.
- Vals AI index: Argon placed first at 68.9% on Vals’ blended index and was in the top five on 20 of 22 benchmarks, with strengths noted in finance, law, coding, and security (Vals AI reporting).
- Arena.ai human preference: In Text Arena, Argon (High) scored 1, 525 points and ranked ahead of several competitors; in Code Arena (WebDev) Argon reached 1, 679 points and climbed to 8th, a large improvement from prior Gemini models (Arena.ai results).
Important caveat: these are single-source evaluator results. Each benchmark uses different datasets, scoring rules, and generation parameters, like temperature, max tokens, and chain-of-thought prompts. Where evaluators’ methodologies were not fully published, observed token counts and accuracy figures should be treated as indicative rather than definitive.
Cost example that makes the tradeoff concrete
Token efficiency is the practical cost lever. Artificial Analysis’ reported averages (Argon ≈ 62k output tokens per Intelligence Index task; Astra ≈ 27k) are dramatic. Using the published promotional rates, that evaluator estimated cost per Intelligence Index task as:
- Argon (promo): $1.99 per task.
- GPT‑6 Astra (promo): $3.26 per task.
But after Argon’s promo rates revert to regular pricing, the same task estimate rises to approximately $3.98 per task for Argon, roughly 20% above Astra under the same assumptions. The takeaway: a lower per-token price does not guarantee lower bills if the model returns much longer outputs.
Short illustrative math: at $20 per million output tokens, a 10, 000-token response costs $0.20, and a 60, 000-token response costs $1.20. Multiply that by volume and the difference becomes material.
How to evaluate Argon responsibly, do this first
Before you consider migrating core workflows, run pragmatic, measurable tests. Prioritize a two-week sandbox and instrument token accounting end-to-end.
- POC 1, Cost and token profile: Replay representative prompts with comparable sampling settings (temperature, max_tokens, stop tokens). Capture input and output token counts and compute per-interaction cost at both promo and regular list prices, including cache read and write assumptions.
- POC 2, Long-decode proof-of-concept: Exercise Long Decode Continuation on a real long reasoning task. Validate whether server-side state is preserved, how resumed segments are billed, and whether output consistency holds across resumed segments.
- POC 3, Safety, tooling, and integration: Test agent integrations, tool calls, content filters, and logging. Confirm the guardrail differences between early-access pilots and paid access, and ask for the formal safety documentation Google will require for production customers.
Immediate action items for procurement and engineering teams:
- Request Google’s billing rules for Long Decode Continuation and cache write pricing (listed as not disclosed for Argon).
- Instrument and cap token usage during trials to avoid surprising bills.
- Ask Google for hardware and latency expectations for sustained 1M-token runs; large context windows have nontrivial runtime and memory implications.
Where Argon clearly moves the needle, and where it doesn’t
- Strengths: Argon delivers strong agentic performance in some benchmarks, wins human-preference text tests, and shows low measured hallucination rates on at least one evaluator. Its 1M token limits enable new long-document and multimodal workflows.
- Weaknesses and unknowns: Measured accuracy on some factual benchmarks remains middling. Argon’s higher output verbosity inflates costs and complicates TCO planning. Google’s ecosystem, apps, agent workbench, and enterprise workflow integrations, still lags competing workspace products, and the concrete rollout timeline and guardrail differences between pilot and paid tiers remain unclear.
Open questions you should demand answers to
- How exactly does Long Decode Continuation preserve state and how are resumed calls billed?
- What are cache write charges, TTLs, and deduplication rules for Argon (Google lists cache read discounts but cache write operations are not disclosed)?
- How stable are Argon’s token-usage characteristics under different prompt templates and generation settings, and do defaults produce chain-of-thought verbosity?
- When will paying API customers and Google AI Ultra subscribers get access, and will the promo pricing apply to early paid customers or only to limited pilots?
- Which guardrails will be present for paying customers versus early internal or Fairwind testers?
Decision checklist for executives (yes/no)
- Do you need >100k token context or very long single-pass outputs?
If yes, test Argon; if no, the extra scale may not justify switching. - Can you tolerate per-interaction cost variability while you instrument token usage?
If no, delay migration until billing rules are clear. - Do you require strict factual accuracy guarantees for regulatory or safety reasons?
If yes, validate Argon’s accuracy on your domain data before moving critical flows. - Do you depend on mature agent/workspace integrations today?
If yes, compare vendor tooling readiness (Google’s apps currently lag some competitors).
Short Q&A, quick answers to likely questions
- Does Argon actually let you generate a million-token output in a single pass?
Google/DeepMind reports support for up to one million output tokens and an equal input context window. Practical latency, hardware, and runtime limits for live services will depend on Google’s infrastructure and any additional per-request limits they enforce. - Is Argon cheaper than GPT‑6 Astra for long tasks?
On the Artificial Analysis Intelligence Index task, Argon’s promo pricing produced a lower estimated per-task cost ($1.99 vs $3.26 for Astra). However, because Argon produces more output tokens per task, that advantage shrinks or reverses when promotional pricing ends (Artificial Analysis’ estimate: Argon ≈ $3.98 per task post-promo, roughly 20% above Astra under the same assumptions). - Are Argon’s answers more factual?
Artificial Analysis reported a relatively low hallucination rate (15%) on its AA‑Omniscience benchmark but measured Argon’s accuracy at 50% on that test family. Human preference tests rate Argon highly for text style, so you may get better prose even when factual accuracy is mixed, verify against your domain ground truth. - Can I rely on Long Decode Continuation for production workflows?
It’s promising for long generations, but Google has not fully documented how state is preserved, how resumed segments are billed, or the latency tradeoffs. Get those details and run a long-decode POC before production adoption. - Should I replace my current model provider with Argon today?
Not by default. Argon is worth piloting if you need its scale or agentic improvements, but don’t switch without POCs that measure token consumption, accuracy on your data, tooling readiness, and a clear billing contract.
Final perspective, what matters most
Argon repositions Google back in the frontier conversation. It brings genuine capability increases, especially for long context use cases and some agentic tasks, and its promotional pricing looks disruptive on paper. Real costs and operational behavior depend on token efficiency, Long Decode Continuation billing, caching rules, and Google’s enterprise integrations. The model that wins in production won’t be the one with the loudest benchmark score. It will be the one with predictable economics, clear runtime semantics for long decodes, and the integrations and guardrails that enterprise teams need. Test, measure, and insist on those assurances before you make a strategic move.