AI agents do more of the work in model development, but humans still make the decisions
Executive summary
- Agent activity rose sharply: the median agent-actions-to-human-inputs ratio climbed from 11 to 28.5 over four weeks, the metric counts discrete agent-initiated steps versus explicit human inputs, as reported by the Atria team.
- Scale without surrender: agents proposed many methods and executed most action steps, but humans retained control of goals, scope, and most final choices (humans made 85.5% of method/parameter decisions and 93.4% of final goal/scope decisions, per the report).
- AI enabled work humans wouldn’t otherwise attempt: of 455 completed AI-assisted tasks analyzed, 151 (≈33%) were judged infeasible without AI.
- Governance gap: long chains of agent activity create a “rubber‑stamp” risk where human review becomes superficial unless oversight is redesigned.
What the team measured
The Atria project (including researchers affiliated with Fudan University) built an agentic language model called Atria Dawn Preview, a mixture-of-experts architecture with 744 billion parameters, and instrumented its development pipeline to log per-step actor attribution (timestamps and actor IDs for each proposal, execution command, and approval). Jonathan Kemper reported the results for The Decoder on September 27, 2026.
The instrumentation tied tasks to execution environments so outputs could be validated against tests or metrics. The report’s logs are reported as 769 task entries in the schema while the article text refers to “more than 700 task logs”; the project analyzed activity from 56 participants. The report does not provide a full public accounting of any task exclusions, but the numbers used in the analysis include 455 completed AI-assisted tasks and 588 tasks with recorded difficulty ratings (see breakdown below).
Clear numbers you can use (and what they mean)
- Logged tasks and participants: 769 task logs are listed in the report schema; the article text describes that as “more than 700” task logs. Analysis covered work from 56 participants.
- Completed AI-assisted tasks: 455 tasks were analyzed as completed with AI assistance (the report does not enumerate exclusions that reduced 769 logged tasks to the 455 completed tasks used for this specific completion analysis).
- Agent usage: agents were involved in 96.5% of the tasks reviewed (as reported).
- Agent activity growth: median agent-actions-to-human-inputs rose from 11 to 28.5 over four weeks (this ratio measures discrete agent actions per explicit human input).
- Work enabled by AI: of the 455 completed AI-assisted tasks, 151 (≈33.2%) were judged infeasible without AI; these 151 tasks were spread across 27 of the 56 participants.
- Decision patterns, methods & parameters: the most common pattern was “AI proposes, human selects, ” observed in 55.4% of method/parameter decision instances (reported by the team). Overall, humans made 85.5% of decisions about methods and parameters while AI made 9.2% (the report does not publish the raw n for this specific breakup; these percentages are the team’s reported figures).
- Goals and scope: according to the report, humans made the final decision on goals and scope in 93.4% of cases. Even for the 151 tasks judged infeasible without AI, humans chose the goal in 95.4% of those instances.
- Task difficulty and resolution (n=588 tasks with difficulty recorded):
- 76% progressed only after human intervention (this 76% is a top-level category that includes smaller subcategories below).
- 23% were solved by the agent without further human action.
- Within the human-intervention category, 3.2% received partial edits from humans and 0.7% were fully taken over by humans (these subcategories are included inside the 76% human-intervention total).
- Iterative correction: when AI outputs required revision and humans provided feedback, the agent executed the subsequent changes on its own 75.4% of the time (reported; this describes agent responsiveness after human-guided correction).
- Benchmarks: the team reports Atria leading on 5 of 16 benchmarks, with named strengths including AutomationBench, CyberGym, and MLE-Bench Lite; the model trailed on GDPval and SWE-Bench Pro in the reported comparisons.
How to read these results: more activity ≠ more autonomy
There’s a crucial distinction. Higher counts of agent actions do not mean agents are making higher-level decisions. The rise in agent-actions-to-human-inputs mostly reflects humans designing workflows that let agents run longer chains of work, not agents independently setting goals or research directions.
The dominant decision pattern, “AI proposes, human selects, ” explains this. Agents generate options and execute steps, but humans still decide what to pursue and where to stop. That keeps strategic responsibility with people while shifting execution to machines.
The rubber‑stamp risk (and why leaders should care)
Long automated chains create a scaling problem for oversight. When a single human signoff unlocks hundreds of automated steps, it becomes impractical for reviewers to inspect every intermediate result. The Atria team calls this the “rubber‑stamp” risk: oversight that becomes cursory because humans cannot audit the volume or complexity of agent-generated intermediates.
Operational behavior in the study amplified the problem: participants sometimes ran agents in “autonomous mode” (letting agents continue without step-by-step approval) to avoid interrupting long runs. That convenience improves throughput but raises the chance that errors, shortcuts, or poorly scoped experiments escape meaningful review.
Industry signals echo the concern: researchers, employee groups, and independent studies have raised alarms about automating higher-order research decisions. Other work, for example, a study co-authored by researchers at Princeton and the UK AI Security Institute, finds that frontier models can perform substantial research engineering tasks but still stumble on the judgment calls that determine research value. These are reported positions and should be interpreted as context rather than proof of inevitability.
Practical guardrails that scale with agent activity
Here are concrete, implementable controls that align with the study’s findings and reduce the rubber‑stamp risk.
- Define decision boundaries: require explicit human signoff for goals and scope. Allow agents to pick methods/parameters only within pre-authorized envelopes and documented ranges.
- Instrument and surface metrics: track an “agent-actions-to-human-inputs” metric per project. Alert on sudden jumps (example threshold: a doubling within a week) and investigate workflow changes that caused it.
- Make execution environments auditable: log inputs, intermediate artifacts, test outcomes, and actor IDs so reviewers can reproduce and validate results without replaying every action.
- Sample-check long runs: instead of exhaustive review, combine randomized sampling of intermediate states with automated correctness tests that run in trusted execution environments. Example policy: require at least 3 random checkpoints per long run plus end-to-end automated tests before final signoff.
- Limit autonomous runs by risk tier: set policy-based thresholds. Example guardrail: permit unattended autonomous runs for low-risk, bounded experiments under 24 hours and with ≥X passing tests, but require human checkpointing for runs that modify production systems, exceed 24 hours, or touch sensitive assets.
- Embed human-in-the-loop tooling: design interfaces that present agent proposals as discrete, reviewable items and require granular approvals for actions that change code, data, or model training regimes.
Open questions executives should watch
- Will agent activity become genuine autonomy? The Atria study stops short of that claim: increased activity so far appears to be chiefly workflow-level, humans still set goals and make final judgments.
- Can models reliably propose and evaluate research directions? Agents can enumerate options and run experiments, but reliably valuing a research direction before experimental data exists remains unresolved.
- What’s the right governance model for production vs. experimentation? Treat experimentation as a place to push autonomy with tight logs, while reserving production changes for stricter human checkpoints and auditable approvals.
Immediate next steps for leaders
- Instrument current pipelines: add per-step actor attribution, execution-environment logging, and the agent-actions-to-human-inputs metric.
- Define who signs goals: require named human owners for goals and scope before agents run at scale.
- Create tiered autonomy policies: categorize runs by impact and apply stricter checkpointing to higher-risk tiers.
- Pilot sampling and automated checks: run a sampling + automated-test regime on a safety-critical project within 30-60 days to validate oversight at scale.
Key takeaways, questions executives ask
-
Who ran the study and where was it reported?
The work was produced by the Atria team (including researchers affiliated with Fudan University) and reported by Jonathan Kemper for The Decoder on September 27, 2026.
-
How much of the work did AI do versus humans?
The report shows agents performed a growing share of action‑level work: median agent-actions-to-human-inputs rose from 11 to 28.5 over four weeks. Despite that, humans made 85.5% of decisions about methods/parameters and 93.4% of final decisions on goals and scope (these percentages are reported by the team; the underlying raw counts for some breakouts are not fully enumerated in the public summary).
-
Did AI enable work humans wouldn’t attempt otherwise?
Yes. Of the 455 completed AI-assisted tasks analyzed, 151 (≈33.2%) were judged infeasible without AI; those tasks were distributed across 27 of the 56 participants.
-
What governance risk should executives prioritize?
The “rubber‑stamp risk”: long chains of agent work can outstrip human review capacity, turning oversight into cursory approval unless organizations implement sampling, automated checks, and clear human checkpoints for goals and scope.
Agents already shift much of the heavy lifting in development workflows. The pragmatic move for leaders is not to resist automation but to redesign decision processes so execution is delegated while strategic control and meaningful oversight remain human responsibilities, enforced by metrics, logs, sampling, and clear policies that scale with agent activity.