The model moat is shrinking. The durable lead is the system that runs the model.
Recent months have confirmed what many executives suspected: standalone model weights, the raw checkpoints you can download or call over an API, are easier to approximate than before. Public benchmark leaderboards and a wave of lab reports show open‑weights from several non‑U.S. labs scoring close to frontier levels on many standard tasks. That does not mean every Chinese or open model matches every U.S. model on every axis, but it does change the commercial calculus.
Why does this matter? Once capabilities are visible through outputs or released weights, competitors can copy, repackage, and re‑weight those behaviors quickly. The defensible, long‑lasting advantage is moving up the stack: owning compute and hardware provenance, the agent orchestration, and the running operation, the telemetry loop that learns from real deployments. Call it system sovereignty: control of silicon, runtime, and the customer feedback loop.
What the public evidence actually shows
Benchmarks and public evaluations now paint a messy picture rather than a single ranking. Different tests emphasize different things: single‑shot competence, repeatability across runs, long‑running agentic tool use, or behavior under safety filtering. Which model “wins” depends a lot on the metric and who ran the test.
- Independent benchmarking projects such as Artificial Analysis run multiple agentic and capability suites that emphasize multi‑step workflows and tool use, and they make clear that agentic evaluation changes the leaderboard dynamics.
- Cyber‑capability evaluations and guarded internal tests show a split between raw capability and safely deployed behavior. Some labs report high capability in controlled settings, while their public siblings, intentionally filtered, show much lower surface capability.
- Single best‑of‑N runs can look impressive. Reliability metrics, for example pass‑style measures that require repeated success, often tell a different story about repeatable production readiness.
Bottom line for buyers: look past headline leaderboard positions. Ask which benchmarks were run, whether results are single best runs or averaged and repeatable, and whether safety filters affected the outcome. Labs and benchmarkers use different methods, so provenance matters as much as the number.
Distillation: a plausible accelerator, but not the whole story
How did several non‑U.S. open models close ground so fast? Two mechanisms are both plausible contributors.
- Output‑based distillation at scale. Policymakers and some labs have warned that large, coordinated API extraction campaigns can be used as “teachers” to train new models. The U.S. Office of Science and Technology Policy issued a National Security and Technology Memorandum (NSTM) in April 2026 that explicitly flagged adversarial distillation as a concern. Reporting that summarizes lab allegations (for example, technical writeups compiled by independent observers) describes campaigns that routed millions of queries through many accounts to capture chain‑of‑thought, tool calls, and safety boundary behavior. Those writeups provide strong circumstantial indicators, though public forensic traces tying a specific student training set to a named teacher model remain limited.
- Native engineering and post‑training pipelines. Domestic R&D, improved pretraining data, better alignment of reward models, and aggressive post‑training engineering, safety tuning, tool integrations, and fine‑tuning on domain data, all plausibly explain rapid improvements. Lab reports and public release notes show iterative gains independent of any single allegation of copying.
Put another way, the evidence supports both copying and genuine homegrown progress. The policy implication is the same either way: if capability is observable, it spreads fast. That makes the surrounding system the strategic prize.
Why the system, not the weights, is the new moat
Three components are now the decisive sources of commercial defensibility.
- Compute and hardware provenance. Access to dense accelerator clusters, high‑bandwidth memory, and reliable regional capacity shapes cost, latency, and the ability to run long agent sessions. Analysts and policy papers emphasize that semiconductor supply and deployment scale are central to how quickly adversaries can operate at scale.
- Agent harness and orchestration. A model inside a well‑designed harness, with retrieval layers, tool routing, state management, context prefilling, and sandboxing, behaves very differently from the same model exposed as a plain API. Benchmakers and labs report significant performance swings from harness changes. More state, better retrieval, and smarter tool orchestration often deliver more value than incremental improvements to base model weights.
- Running operation and the deployment feedback loop. Companies that embed models into real business workflows collect telemetry: failure modes, drift signals, tool latency, downstream user corrections, and business outcome metrics. That operational data, and the teams that act on it, embedded engineers, SREs, and product owners, are the levers that turn capability into customer value.
These are glue assets: expensive to build, slow to copy, and tightly coupled to customer relationships. They are where economic moats form now.
Security and the capability/safety tradeoff
Evaluations that stress offensive cyber capability or unconstrained code execution reveal a tension. Labs that expose raw capability in controlled settings report higher scores on offensive‑style benchmarks. Their publicly available versions, with upstream safeguards, show far less surface capability. That delta matters for enterprise buyers. Do you want raw capability you must police, or a constrained product you can govern easily?
Regulators and lab operators are reacting. The OSTP memorandum and follow‑on policy discussions recommend export controls, access gating for high‑capability offerings, and more rigorous audit mechanisms. Industry groups are pushing back against broad restrictions that could stifle research and legitimate commercial activity. The political balance is unsettled, and no policy can make observed outputs un‑observable. Policy can only raise the cost and friction of large‑scale extraction.
Commercial plays that buy defensibility
Expect three practical, repeatable business strategies to dominate:
- DeployCo‑style services. Embed engineers with customers, instrument production workflows, and sell continuous operational improvement, not just API access.
- Chip + cloud integration. Offer tightly coupled hardware, runtime, and orchestration so customers get predictable latency, proven sandboxing, and context memory tuned for agents.
- Certified, controlled access. Provide gated programs that let critical customers use more powerful models under audit, specialized contracts, and accountability mechanisms.
These approaches trade one‑time licensing for ongoing value capture: SLAs, telemetry sharing, and a subscription to operational competence.
What executives should do now
Prioritize system questions, not just model names. Practical next steps:
- Inventory and risk‑rate current AI use. Map where models run, what telemetry exists, and which vendor relationships include deployment engineering or simply an API key.
- Buy or build a harness. Get basic retrieval, tool controls, context management, and observability into place before you scale agentic use.
- Negotiate telemetry and training clauses. Require vendors to disclose whether they train on customer data, and if they do, insist on opt‑in, audit logs, and contractual limits on reuse.
- Secure compute provenance. Know where your cloud provider sources accelerators and whether you can obtain geographic or contractual guarantees of capacity.
For many companies that means shifting budget from model licensing toward deployment engineering, observability, and vendor SLAs. A vendor’s leaderboard trophy is useful marketing, but your legal exposure and production reliability depend on the harness and the telemetry pipeline.
Vendor RFP questions (copyable)
- Describe your deployment feedback pipeline: what telemetry you collect, how quickly you ship model fixes, and what SLAs you offer for drift mitigation?
- Do you train on customer data by default? If so, what opt‑in/opt‑out controls, audit logs, and data‑provenance measures do you provide?
- Explain hardware provenance and continuity: where are accelerators sourced, what redundancy and regional guarantees exist, and what contractual remedies cover supply disruptions?
Suggested timeline for leaders:
- 0-3 months: inventory, vendor RFPs, and a pilot harness for one high‑value workflow.
- 3-9 months: instrument production, negotiate telemetry training terms, and integrate basic tool controls and sandboxing.
- 9-18 months: operationalize continuous improvement, embed or contract deployment engineers, and codify SLAs and audit practices.
Policy and Europe: what to watch
Governments are treating distillation as a national security and commercial policy issue. The OSTP memorandum made the risk visible at the U.S. policy level, and think‑tank and academic memos argue for a blend of semiconductor controls, access gating, and accountability for model providers. Those levers matter. They raise the cost of large‑scale distillation, but they do not eliminate the core technical problem that outputs are observable by design.
Europe faces a particular strategic choice. Building the full stack, regional compute, certified hosting, and an onshore deployment knowledge base, is the only way to preserve operational sovereignty. Several European firms and initiatives are actively trying to assemble pieces of that stack, but creating ample regional capacity and the operational know‑how will take time and capital, and time is the scarce resource.
A short counterpoint
Do not read this as fatalism about weights. Algorithmic breakthroughs, new learning paradigms, and community innovation could shift the balance again. More efficient architectures, improvements in continual learning, or open research advances might re‑separate capability from deployment. The current trend, though, favors players who can integrate models into reliable, observable systems at scale.
Key takeaways, questions you should be asking
-
Has China actually closed the gap?
Public leaderboards and lab reports show several open‑weights approaching frontier performance on many public tests; exact margins vary by benchmark and by whether safeguards were applied. The trend is clear: the observable gap has closed the gap in many public metrics. -
Is distillation the main cause?
There is substantial circumstantial evidence, including policy notices and technical writeups summarizing alleged large extraction campaigns, that output‑based distillation accelerated catch‑up. Conclusive, public forensic links tying a specific student dataset to a particular teacher model are less common; both copying and native engineering likely played roles. -
Can export controls stop it?
Export controls and supply‑chain enforcement raise costs and complicate large‑scale attacks, but they cannot prevent an attacker who can observe outputs from copying behavior entirely. Policy must be paired with technical and commercial measures. -
Where should durable advantage be built?
In the system: regional compute and hardware guarantees, robust agent harnesses with retrieval/tool controls, and the running operation that collects and acts on deployment telemetry. Those things are slow to build and hard to monetize via a leaked checkpoint. -
What should Europe do?
Prioritize system sovereignty: invest in regional compute pools, certified hosting for critical workloads, and incentives for deployment partnerships that keep operational knowledge onshore. Time and capital allocation are the strategic constraints.
Reported industry moves reflect the same reality: firms are shifting from selling checkpoints to selling continuous operational capability and audited access, a commercial recognition that the product is the running operation.
The practical lesson is straightforward for leaders: stop evaluating vendors on a single leaderboard metric and start auditing their harness, compute provenance, telemetry pipeline, and contractual commitments around customer data and model updates. Weights matter, but the business you buy is the system that runs the weights in production, and that system is where value will be captured for the foreseeable future.