Sept. 12-13, 2023: How a staged safety plan became a debate about incentives
On Sept. 12 Anthropic CEO Dario Amodei published an essay titled “We Must Pace the Frontier, ” which proposed a staged approach to slow how quickly frontier AI capabilities improve while expanding independent safety testing and oversight. The next day Solana co‑founder Anatoly Yakovenko posted on X: “Profitability at $1 trillion mcap.” He did not name a company or provide evidence for that claim.
Public reactions came quickly. OpenAI CEO Sam Altman posted on X that he agreed with Amodei and said, “Committing to having independent evaluators with employee‑like access is a great idea, and we will do the same.” Reuters and the Financial Times reported Elon Musk reposted Amodei’s essay and commented, “Dario is right.” David Sacks pushed back on X, arguing frontier labs could slow development voluntarily and warning against regulatory carve‑outs.
What Amodei actually proposed
Amodei’s framework spells out stages and specifics. It is not a call to stop research:
- Embedded independent evaluators: third‑party reviewers with access comparable to internal risk staff, Amodei named METR (Model Evaluation and Threat Research) as an example, able to inspect models and safety practices and to publish findings subject to narrow redactions.
- Coordination among frontier labs in democracies: common safety requirements and deliberate limits on unchecked capability growth, while recognizing potential antitrust complications and the possible need for narrow government authorization.
- International coordination: negotiated agreements for narrow, high‑risk areas (for example, AI‑assisted biological threats and rapid automated model improvement), with Amodei noting a broad comprehensive pause is unlikely because verification would be extremely difficult.
Why the four words mattered, and what they didn’t prove
Yakovenko’s post hit on a sensible point: proposals about safety can shift incentives and market narratives. That’s a valid governance concern. But the line “Profitability at $1 trillion mcap.” came without support, no company named, no valuation method offered, and private IPO chatter is not the same as a public market capitalization.
Reporting said Anthropic and OpenAI were preparing for potential IPOs, but both remained private at the time. Any dollar figure tied to them is an estimate, not a verified public market cap. Treating private secondary valuations or IPO talk as proof of a $1 trillion target mixes speculation with fact.
The practical knots executives need to see clearly
Amodei’s stages are plausible policy. Turning them into practice, though, creates real legal, security, and commercial trade‑offs. Three stand out.
- Embedded evaluators: access vs. legal constraints. Anthropic described reviewers having company‑laptop access, office access, permissions like internal risk staff, and the ability to publish findings with limited redactions. It is not simple. IP claims, export‑control rules, privacy laws, and national‑security obligations can legally block sharing certain training data, model weights, or logs. A typical clash: a reviewer asks for datasets that contain third‑party personal data or export‑controlled biological sequences, and the company cannot legally transfer the raw files.
- Coordination and antitrust exposure. Narrow safety collaboration differs from competitor price‑fixing, but the line can be thin. How meetings are recorded, who attends, and whether safety requirements are enforced by the companies or by an external body all affect antitrust risk. Companies should assume scrutiny and design coordination with neutral facilitators, clear charters, and, where appropriate, early consultation with antitrust counsel or a DOJ/FTC advisory process.
- International verification is hard. Software and model improvements are easy to copy or reimplement; capabilities can be hidden inside derivative work. Amodei acknowledged that a broad pause would be difficult to verify. International agreements are therefore better suited to narrow, high‑consequence risks than to broad capability freezes.
Cybersecurity and the downstream risk clock
Operational risk rises as capabilities accelerate. A Bank for International Settlements paper warned that advanced AI could shorten software‑patching windows, from weeks to minutes, because automated tools may discover and chain exploits far faster. This is not theoretical for C‑suite resilience planning: faster exploit cycles require faster monitoring, shorter incident response SLAs, and tighter devsecops integration.
Three immediate actions for executives
- Inventory your frontier exposure in 30 days. Identify which models, suppliers, and integrations you rely on that qualify as “frontier”, the highest‑capability models or services that materially change risk profiles.
- Run a 90‑day external‑review pilot. Contract a neutral reviewer (academic lab, NGO, or independent firm), define access levels, run two redaction scenarios, and document results and legal constraints. Use the pilot to produce a repeatable playbook.
- Engage legal and regulators early. Consult antitrust counsel and, where relevant, the DOJ Antitrust Division or equivalent authorities before any cross‑lab coordination. If export controls or privacy laws constrain sharing, document the legal basis and propose technical mitigations (e.g., synthetic datasets, aggregated telemetry) in your review charter.
Operational checklist, what to build now
- Define “embedded evaluator” in contract language: access scope, physical vs. remote access, data handling, publication rights, and approved redaction categories.
- Draft NDAs and limited‑purpose data‑use agreements tied to reviewer charters; include liability allocation and indemnities, and obtain an insurance rider where possible.
- Create an escalation protocol and whistleblower channel routed through independent counsel and governed by a standing third‑party arbiter for disputed redactions or findings.
- Document safety coordination as safety‑only, facilitated by a neutral third party, with minutes and a public charter to reduce antitrust risk.
- Stress‑test incident response timelines to account for compressed exploit windows; shorten SLAs for monitoring, patching, and revocation where frontier models are in production.
Where public commitments stop and detail starts
Committed language matters: Amodei set out the stages and named METR as an example; Altman publicly committed OpenAI to embedded independent evaluators. But those commitments were preliminary. At the time of the announcements, neither Anthropic nor OpenAI had published contracts, named contracted evaluators, provided start dates, or released full governance documents. That gap invites skepticism and competitor signaling.
“We must slow the pace at which we improve the capabilities of AI models.”, Dario Amodei
“I agree with Dario that we need to pace the frontier.”, Sam Altman (X)
“Profitability at $1 trillion mcap.”, Anatoly Yakovenko (X)
Bottom line for the C‑suite
Support for independent safety review is something executives should welcome. Endorsing a principle is not the same as operationalizing it. The real work is designing access rules that respect IP, export controls, and privacy; documenting coordination to limit antitrust exposure; and proving to investors and regulators that review mechanisms are real, repeatable, and independent. Treat governance like product engineering: specify requirements, prototype the solution, measure compliance, and iterate.
Key questions leaders are asking, short evidence‑based answers
- Did Amodei call for stopping AI research?
No. Amodei proposed “pacing” capability improvements through staged oversight and more testing, not an outright halt to research.
- Did Altman and Musk publicly endorse parts of the plan?
Yes. Sam Altman publicly agreed with Amodei and committed to embedded independent evaluators; Reuters and the Financial Times reported Elon Musk reposted Amodei’s essay and commented “Dario is right.”
- What did Yakovenko’s “$1 trillion mcap” claim prove?
It signaled skepticism about motives but supplied no evidence, no company name, and no valuation calculation to substantiate the allegation.
- Are embedded evaluators already operating?
Companies described intentions and general access levels for outside reviewers, but public contract details, start dates, and governance documents were not released alongside the initial statements.
- Will international agreements stop risky AI development?
Amodei himself flagged verification challenges and treated a broad pause as unlikely. International coordination on narrow, high‑risk use cases may be feasible, but enforcement and verification remain difficult.