Executive summary
This week’s signal: OpenAI published GPT‑6 (Sol and Luna) with illustrative pricing examples and bundled developer tooling (GPT‑Live‑1 / Agents API). According to OpenAI, the combination of lower per‑token cost and live/agent features changes the economics and engineering effort for multi‑step autonomous workflows, these are vendor‑reported results and should be validated against your workloads and independent trackers.
The headlines (and the exact figures to know)
-
OpenAI, GPT‑6 (Sol & Luna): According to OpenAI’s announcement, they show illustrative price examples that include reductions such as $20 → $10, $0.20 → $0.10, and $1.20 → $0.50 per 1M tokens. OpenAI also reports benchmark comparisons where, for example, GPT‑6 Sol (xhigh) scored 33.2% on the composite OpenAI published and a reported cost per task of $0.27, versus Claude Opus 5 (max) at 26.9% and a cost‑per‑task OpenAI reports as roughly 11.1× higher on that same composite. These numbers are on OpenAI’s page and are vendor‑reported; see OpenAI’s announcement for the full tables and methodology.
OpenAI, Introducing GPT‑6: Sol & Luna - OpenAI, GPT‑Live‑1 & Agents API: The developer guide documents realtime audio/voice, persistent sessions, tool/plugin integrations, vaulted secrets, sandboxes, and observability features useful for building AI agents. These reduce some engineering lift for live, agentic experiences while increasing operational and security responsibilities. OpenAI Developers, GPT‑Live‑1 & Agents API
- Anthropic released Claude Opus 5.5.
- xAI (Grok) announced Grok 4.7.
- Google (Gemini) published Gemini 3.8 Live Avatar and a Flash TTS variant for lower‑latency speech.
- Microsoft updated Copilot with a refreshed experience described as “home, code, and autopilot.” Microsoft, New Copilot
- Spotify adjusted how listener Taste Profiles shape home feeds.
- Meta had its Connect 2026 announcements. Meta Connect 2026
- Google Research published Project Suncatcher.
What the benchmark numbers actually mean
Two caveats: (1) the score percentages and cost‑per‑task figures published by OpenAI are vendor‑reported and use the composite benchmark mix OpenAI selected, and (2) independent trackers can show different rankings depending on which tests you weight.
OpenAI references a set of benchmark suites; here’s a one‑line primer on each so you can map them to your priorities:
- AutomationBench, tool‑driven, multi‑step workflows across business apps (vendor links include third‑party hosts such as Zapier).
- Agents’ Last Exam, long‑horizon professional workflows spanning many sub‑industries.
- FrontierCode / DeepSWE, coding quality and merge‑readiness (tests on real codebases and developer tasks).
- OSWorld, long‑horizon offline computer‑use tasks and simulated user interactions.
- Others, OpenAI references additional third‑party and internal tests;
Further viewing and reading
If you want a quick visual roundup and a deeper comparative analysis, these two resources are handy.
- Video roundup: Opus 5.5, GPT‑6 Sol, Jev & Muse
- Comparative analysis: ChatGPT vs Claude vs Gemini vs Grok