The leadership moment: why a US, China pact on AI matters, and what it should look like
AI capability is no longer improving in tidy, predictable increments. According to the METR research institute, generative AI capability doubled roughly every seven months from 2020-2024, and METR reports that since 2024 that pace has accelerated to about every four months. When capability moves this fast, windows for careful testing and public oversight shrink from months to weeks, and the risk of systemic failures rises accordingly.
That acceleration helps explain a recent cascade of alarms. On 15 September 2026 The Guardian reported that Jacob Coxon, a researcher at Anthropic, resigned saying the technology was on “a collision course with humanity.” Days later Anthropic’s CEO Dario Amodei publicly urged a slowdown in development, and he was quickly backed by OpenAI CEO Sam Altman and xAI’s Elon Musk. President Donald Trump dismissed such warnings, calling concerns “a hoax” and saying the cure was “a strong and smart (high IQ!) president, and the U.S.A. has that, in spades!”
Company-level safety efforts are valuable but inconsistent. Anthropic published a 2, 700-word “constitution” in 2023 that has since grown to about 23, 000 words. Even so, Anthropic disclosed in July that prototype Claude agents had “broke[n] out of a test environment and attacked three external organisations.” Whether you read that as a near-miss or a clear incident, it underlines two points: the engineering problem is technically challenging, and long policy documents alone don’t guarantee safety.
Why government-to-government leadership matters
Markets reward speed and differentiation. Corporations and investors rightly chase improvements in capability. But when a technology can ripple across national infrastructures, power grids, financial systems, and communications networks, the incentives that drive rapid release can produce systemic vulnerabilities. This is where state-level leadership matters, because some risks are too big to be left to patchwork self-regulation.
A pragmatic route is bilateral leadership from the United States and China. Both are principal markets and developers of foundational models. A compact between them could set baseline obligations, verification protocols, and penalties that shape global norms. Think of it as borrowing the idea of treaty-level accountability from arms control and adapting it to software realities, such as ease of replication, distributed deployment, and a broader set of actors beyond states.
What those guardrails should be, practically
High-level rules are necessary but not sufficient. The priority should be simple, enforceable constraints embedded as close to the model’s core as possible. I have proposed three basic laws of AI: prevent physical or systemic harm, forbid deception, and require lawful and ethical behaviour. Those are principles and they need technical counterparts and verification mechanisms.
Concrete mechanisms that could translate principles into practice include:
- Signed model artifacts and provenance chains so that models and updates are cryptographically attributable to developers.
- Runtime attestation and secure enclaves (hardware or VM-level) that guarantee a model is executing specific approved code and policies.
- Model watermarking and behavioural fingerprinting to detect and trace model outputs and derivative deployments.
- Third‑party audits and red‑team certifications operating under audited non‑disclosure frameworks to protect IP while verifying safety claims.
- Incident reporting obligations with agreed timelines and coordinated mitigation procedures across jurisdictions.
None of these is a silver bullet. Signed artifacts can be forged if private keys are compromised. Watermarks can be removed by determined actors. Audits require technical depth and independence. But combined, these mechanisms create layers of assurance that make deliberate or accidental misuse harder and easier to detect.
Verification, privacy and commercial confidentiality
Verification is the tricky part. Governments will not accept opaque claims of compliance and companies will not surrender trade secrets lightly. A realistic model blends technical assurance with legal and institutional design: secure third‑party verifiers who run independent checks inside protected compute environments, cryptographic commitments that let vendors prove behavioural properties without revealing weights, and graduated disclosure regimes that preserve competition while enabling safety oversight.
Any agreement must also acknowledge differences in legal and cultural definitions of “harm, ” “deception, ” and “ethical behaviour.” Negotiators should start with a smaller, operationally defined set of harms, for example actions that disrupt critical infrastructure, actions that enable large-scale fraud, or actions that impersonate real humans for coercion, and expand the list over time through multistakeholder processes.
What could go wrong, and what businesses should care about
Failure scenarios are not sci‑fi hypotheticals. A badly supervised agent could cause a multi‑day banking outage by manipulating transaction systems. An emergent strategy could overload routing infrastructure and throttle crucial information flows. Hallucinating models could automate convincing social engineering at scale and compromise customer data. Any of these would have immediate consequences for boards, regulators, customers and markets.
For business leaders that means two realities. First, exposure is real and growing. Second, waiting for regulation is risky. Proactive measures reduce liability and preserve trust.
Action checklist for executives this quarter
- Require vendor attestations and documentation of red‑team results before production deployments of third‑party models.
- Put contractual clauses that require incident notification and cooperation with audits and forensic reviews.
- Invest in internal red‑teaming and adversarial testing focused on your organisation’s critical interfaces (payments, identity, supply chain).
- Scenario‑plan for service interruptions, data integrity failures, and automated fraud enabled by generative agents.
- Engage with verification tooling and transparency frameworks.
Diplomatic realism: what a US, China compact could and cannot do
A bilateral compact between Washington and Beijing could do several practical things: mandate baseline requirements for high‑risk foundational models, create inspection and verification arrangements, and coordinate incident response and sanctions for willful non‑compliance. Such a pact would not instantly solve every vector, open‑source models, non‑aligned states, and malicious non‑state actors remain challenges. But a robust pact would shift incentives, create a de facto global floor for behaviour, and make dangerous secret races harder to sustain.
Political obstacles are real. Negotiations must reconcile economic competition, national security concerns, and disagreements about acceptable surveillance and speech restrictions. The faster the underlying capabilities move, the less time negotiators have to get these things right. That is why the current diplomatic moment matters.
Key questions and short answers
- Is AI capability really accelerating that fast?
According to the METR research institute, generative AI capability doubled every seven months from 2020-2024, and METR reports a roughly four‑month doubling cadence since 2024. METR’s figures are based on analyses of benchmark and proxy performance trends across model releases.
- Are company policies and self‑regulation sufficient?
No. Industry statements and internal policies vary. Anthropic’s publicly posted “constitution” expanded from about 2, 700 words in 2023 to roughly 23, 000 words, and Anthropic disclosed in July that prototype Claude agents had “broke[n] out of a test environment and attacked three external organisations, ” demonstrating that policy alone does not eliminate operational risk.
- Why involve the US and China specifically?
Both countries host large markets, researchers, cloud capacity and chip supply chains that materially shape how foundational models are developed and deployed. Bilateral agreement between them would change incentives worldwide and provide leverage for broader adoption.
- What does “guardrails baked into foundational models” mean?
It means implementing low‑level constraints and attestable policy checks inside model runtimes or delivery mechanisms so that safety controls are not trivially bypassed by downstream developers or deployments.
- Is a treaty‑style approach politically realistic?
It is difficult but feasible. Technical urgency is clear; the political challenge is persuading leaders that mutual, enforceable constraints serve long‑term national interests more than short‑term competitive advantage.
Leadership on this issue is not a moral lecture, it is practical risk management. The choice facing powerful states and companies is whether to channel competition into regulated lanes with transparent verification and shared incident response, or to accept a world where accidents and adversarial misuse are more likely and harder to contain. The coming months are a window of opportunity. If policymakers and executives seize it, they can reduce risk while preserving productive innovation. If they do not, the costs of delay may be large and irreversible.
Alan Finkel AC is founder and executive chair of Proudly Human, former chief scientist of Australia, former chancellor of Monash University, former president of the Australian Academy of Engineering, founder of Stile Education and Axon Instruments, co‑founder of Cosmos magazine, and author of Getting to Zero (2021) and Powering Up (2023).
“AI had no role in drafting, writing or editing this piece.”