When a lab tech turns a 48-hour data-cleaning slog into an afternoon with an LLM, it feels like magic. But every hour saved is also an hour available to start another project. A recent theoretical economics paper (arXiv:2607.17397), authored by researchers at Princeton, the University of Washington and other institutions, models that simple arithmetic and reaches a sobering conclusion: if AI mostly speeds the wrong parts of research, we could get more papers that are, on average, less thorough.
How the model works, and the assumptions that matter
The paper adapts optimal foraging theory from behavioral ecology to how researchers allocate limited time across projects. In the setup, each project has a quick viability check and, if continued, a mandatory completion phase plus an optional voluntary deep-dive.
The authors deliberately idealize LLMs so they can isolate the pure time allocation mechanism. Models are treated as effectively error-free and nearly costless time savers. Researchers are modeled as reallocating their saved time to maximize productive outputs under opportunity-cost constraints. Those simplifying assumptions matter, in the real world, hallucinations, validation costs, and mixed incentives will change the arithmetic. But isolating the time effect reveals a clear channel by which AI can reshape scientific effort.
To make the model concrete, map its generic phases to routine research tasks:
- Viability check: quick literature scans, a small pilot analysis, or a test script to see if data exist.
- Mandatory completion: producing figures, basic baseline analyses, documentation, formatting, and submission materials.
- Voluntary deep-dive: extra experiments, robustness checks, preregistered extensions, replication attempts, long-form writing and polishing.
Three channels, three very different outcomes
Where AI reduces time cost matters. The paper lays out three qualitative scenarios:
- AI speeds early idea screening. Researchers become pickier about which ideas to try, but they also start more projects. Example: a group that once ran two full replications might now spin up five quick pilots, increasing starts but reducing robustness per finished paper.
- AI speeds the publish-ready pipeline (writing/formatting/basic analysis). More projects clear the final hurdles and get submitted. Example: boilerplate generation and automated figure writing push marginally finished work into circulation, increasing volume but lowering average thoroughness.
- AI speeds the voluntary deep-dive. Extra time is actually invested into robustness: more experiments, deeper analyses and higher per-paper quality.
In the model’s framework, holding its assumptions fixed, two of these three channels reduce per-project thoroughness because saved time is reallocated to starting new projects rather than to doing deeper work. As the paper puts it: “As a labor‑augmenting technology, LLMs increase the opportunity cost of our time, impelling us to do more, less well, rather than the same amount, better.” The authors therefore label this the “fallacy of saved time”.
“Even if language models worked perfectly, they could make research worse, not better.”, Jonathan Kemper, The Decoder (Aug 23, 2026)
Which real signals line up, and what they actually show
The economics model is theoretical, but several early signals echo its mechanism. These are preliminary, heterogeneous, and must be interpreted cautiously, but they help map the theory onto current experience:
- OpenAI field report: OpenAI summarized eight scientific case studies and reported large speedups in tasks like rewriting research software, in one set of examples described as “up to 60×” speedups, only to find the bottleneck shifted to validation and long-term maintenance.
- METR developer study: In a study of experienced open-source developers using AI tools, measured completion times were reportedly 19% longer even though participants felt 24% faster. That gap between subjective speed and measured time underscores a psychological channel the model sidelines: perceived gains can outpace real productivity gains.
- AI-generated papers and moderation strain: An incident involving Sakana AI’s “AI Scientist-v2”, a mostly AI-generated submission that contained citation errors and reached an ICLR workshop, illustrated both hallucination risks and strained review filtering. Public repositories and preprint servers responded: arXiv tightened moderation and warned of stiffer penalties for hallucinated sources or leftover AI meta-commentary.
These vignettes show two recurring patterns. Big speedups on one task expose new bottlenecks elsewhere, like validation, reviewing, and maintenance. And perceived productivity gains can diverge from objective measures. Remember, the theoretical paper intentionally sets hallucination and cost effects to zero to isolate the opportunity-cost channel. In practice, those omitted terms can amplify or reduce the model’s predicted outcome.
Why incentives and institutions determine whether AI helps or hollows out science
Tools change marginal returns. Incentives determine behavior. If hiring, promotion, and funding systems reward publication count over methodological rigor, researchers who suddenly have more hours will rationally spend them producing more outputs. That creates a collective-action problem: individual gains in throughput can impose external costs on reviewers, replicators, and the scientific record.
Two lessons follow. First, banning LLMs outright is neither realistic nor desirable. Second, institutional design, what we measure, reward, and validate, is the lever that turns time savings into depth or into churn.
Actionable levers, concrete pilots departments, journals, and funders can run in 6-12 months
Practical responses should be discipline-sensitive (the paper recommends tailored policies because which stage AI speeds varies by field), but several pilots can be rolled out quickly to test what works:
- AI-use disclosure pilot (journal level). Require prompts, model version, and a 200-word description of the model’s role for the next 100 submissions in a participating journal track. Measure reviewer time, correction rates, and reproducibility outcomes over 12 months.
- Registered-reports expansion (funders + journals). Funders and a willing journal should run a 12-month pilot that prioritizes a registered-reports track in one field, tracking downstream replication and citation patterns versus a matched control cohort.
- Replication & maintenance line item (grant applications). Require applicants to include a budget line for replication, validation, or code maintenance in grant proposals above threshold amounts, and monitor whether adding this explicit funding changes team behavior.
- Reviewer capacity experiment (publisher level). Recruit additional paid reviewers for a subset of submissions and compare turnaround, reviewer workload, and editorial satisfaction with a control group to see whether extra reviewer capacity preserves quality amid higher submission volumes.
- Reproducibility badge program (department/journal). For a 12-month pilot, flag papers that include runnable code, raw data, and an executable notebook. Track whether these flagged papers require fewer post-publication corrections.
What to measure, concrete signals and where to find them
To test whether LLM adoption is making research more cursory, track a small set of concrete metrics and data sources:
- Submissions per author per year: Crossref/ORCID and publisher APIs.
- Reviewer time per manuscript and reviewer decline rates: publisher internal dashboards or editor surveys.
- Post-publication corrections and retractions: Retraction Watch and journal correction logs.
- Replication success rates: outcomes from registered-report experiments or replication consortia.
- Proportion of papers with runnable code and data: GitHub links, Zenodo deposits, and journal reproducibility checklists.
Key takeaways, quick questions and honest answers
-
Can time-saving LLMs reduce research quality?
Yes, in the theoretical model (arXiv:2607.17397) that isolates time savings, if AI mainly speeds idea screening or the publish-ready steps, saved time is likely to be redeployed into starting more projects, reducing voluntary deep dives and lowering average per-paper thoroughness. -
When do LLMs improve per-paper quality?
Only if they predominantly speed the voluntary deep-dive tasks, extra experiments, robustness checks, and careful analysis, does the model predict higher per-paper quality. -
Do early real-world signals support this mechanism?
Preliminary examples (an OpenAI field report, a METR developer study, and high-profile AI-generated submissions such as Sakana AI’s case) are consistent with the mechanism: speedups often expose downstream bottlenecks and subjective feelings of speed can outstrip measured gains. Still, these reports are mixed and early; rigorous, discipline-specific empirical work is needed. -
What should institutions do now?
Start pragmatic pilots: mandatory AI-use disclosure for a sample of submissions, require reproducibility artifacts, fund replication/maintenance explicitly, expand registered reports, and experiment with paid reviewer capacity. Measure effects and iterate.
A short counterpoint
LLMs can and do cut tedious overhead, like literature triage, boilerplate code, and early debugging. In environments with strong validation infrastructure and aligned incentives, they may free researchers to do deeper work. Smaller teams or those without prior access to large technical support may benefit most. The risk the model highlights is not inevitable; it depends on how institutions react.
A direct call to leaders
Leaders at journals, universities, and funders should treat this as an organizational design problem, not a tech problem. Within the next 12 months, pilot mandatory AI-use disclosure for a bounded set of submissions. Require a replication or maintenance budget line in grant proposals of substantive size. Expand registered reports in one or two fields. Measure reviewer time, post-publication corrections, and replication outcomes. If these pilots show rising churn and falling reproducibility, tighten incentives. If they show more depth, scale the policies. The question is operational: will we let AI be an efficiency hack that boosts throughput, or will we use it to reroute human attention toward the slow, high-value work that actually advances knowledge?