Access to AI makes people far less likely to say “I don’t know”, in low‑information tasks
When people see a quick, confident answer from an AI, they tend to accept it even if it’s wrong. That pattern showed up across five experiments with 3, 132 participants. Simply having AI output available made people much less willing to withhold judgment, raised their confidence, and produced more answers but fewer correct ones.
The experiments used movie visual‑detail questions that rarely appear in text. The authors picked these items because the tested model did poorly on them. That choice matters. This is not sensible delegation to a reliable tool. Participants deferred to an assertive answer even when the AI was usually wrong.
How the experiments were set up
Researchers (Marcoccia, Quattrociocchi, Capraro et al., 2026) ran five studies, labelled 1a, 1b, 2, 3 and 4. They compared three conditions: no AI access, on‑demand AI answers, and automatically displayed AI answers. The AI called in the preprint “Step 3.5 Flash” was intentionally weak on these items. The paper references other work that compares stronger models (e.g., GPT‑5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash) on different question sets.
Some studies added a small financial incentive, ±$0.10 per response (+$0.10 correct, −$0.10 wrong, $0 for “I don’t know”), to see if modest stakes change behavior.
Key results (with study context)
- Overall sample: 3, 132 participants across five experiments.
- Withholding judgment (Studies 1a & 1b): Control groups said “I don’t know” 36% and 44% of the time. With AI access, withholding fell to 6% and 3% respectively.
- Confidence (Study 2): Mean self‑reported confidence rose from 29.6 (no AI) to 75.9 (with AI) on a 0-100 scale.
- Correctness (Study 2): Accuracy dropped from 27.6% (no AI) to 10.0% (AI access).
- Aggregate correctness (no incentives): 27.5% without AI vs. 9.2% with AI.
- Effect of small monetary incentives (Studies 2-4): The ±$0.10 pay structure reduced, but did not eliminate, AI seeking and increased abstention slightly. It failed to restore no‑AI withholding rates.
- AI requests (Study 3): Average AI requests out of 6 possible were 5.27 without incentives vs. 4.53 with incentives.
- Automatic display (Study 4): When AI answers were shown automatically, withholding dropped from ~35% (no AI) to 1% without incentives. With incentives, the fall was from ~39% to 7%.
Why participants deferred: “Epistemia”
The authors call the phenomenon “Epistemia.” People defer to AI outputs because they sound confident and always give an answer, not because they are reliable. Generative models usually return definitive responses instead of “I don’t know.” That makes them seem like constant, assertive advisors, even when a person should abstain.
This result runs counter to classic advice‑taking literature, which finds people usually underweight external advice and move only partially, about one‑third, toward an advisor’s position. Here, simply seeing an AI answer, whether requested or unsolicited, pulled participants’ judgments much more strongly than that historical baseline.
What leaders and product teams should care about
The experimental setting used low‑information trivia where the model was usually wrong. That proves the psychological pull of AI, but it doesn’t tell us how large the effect will be in professional domains where models are generally more accurate. Still, several practical implications are immediate:
- Decision quality can decline while confidence rises. Teams that substitute quick AI answers for verification may make faster but worse choices.
- UI patterns matter. Automatically displayed or prominent AI outputs substantially reduce people’s willingness to abstain. That UI pattern is becoming common in search, writing assistants, and inbox tools.
- Small incentives and warnings aren’t enough. The modest monetary nudge in the experiments reduced AI use slightly but did not restore baseline caution. That suggests stronger accountability, provenance, or reputational incentives may be needed.
Related work echoes these concerns. A Swiss Business School study with 666 respondents found a strong negative correlation between AI use and critical thinking. Microsoft research has warned that generative AI can shift problem‑solving roles and increase demands on metacognition. Both findings reinforce the risk that routine AI access could erode verification habits in organizations.
Practical mitigations you can test now
Design and policy interventions can blunt Epistemia’s pull. Below are practical, testable ideas, with one‑line implementation hints you can act on this quarter.
- Signal uncertainty explicitly. Show calibrated confidence ranges (e.g., “Confidence: 40-60%”) and list the model’s top 2-3 alternative answers so users see that the response is not definitive.
- Add decision friction for low‑confidence outputs. Require an extra click or a short confirmation for AI answers flagged as low confidence. Hide unsolicited suggestions behind a toggle so they must be requested.
- Normalize and preserve “I don’t know.” Make abstention a visible, respected choice. No penalty in scoring or reputation for selecting “I don’t know” in uncertain contexts. Surface examples where abstention was the correct move.
- Record provenance and audit AI assistance. Tag decisions or drafts that used AI and log the model, prompt, and timestamp so reviewers can trace and evaluate AI‑driven choices.
- Train for metacognition. Build a two‑question checklist into workflows for any decision above a threshold: “What would make me change this answer?” and “How could I verify this in under 10 minutes?”
Concrete experiments to run (product & org level)
- A/B test UI patterns: Compare auto‑shown answers vs. on‑demand answers in an internal knowledge tool. Measure abstention rate, correctness (spot‑checked), and downstream rework or error rate over four weeks.
- Stakes experiment: Run domain‑specific trials with higher monetary or reputational incentives (not just ±$0.10) to see whether larger accountability restores withholding.
- Exposure study: Track new users’ abstention and verification behavior over months to detect whether repeated AI exposure produces lasting drops in critical thinking.
Limits and open questions
This research shows a robust effect in a narrow, low‑information task set where the model was intentionally unreliable. Important questions remain. How large and persistent is Epistemia in expert or high‑stakes domains? Can stronger or social incentives, such as reputation or auditing, reverse it? Which UI or training interventions are most effective over the long term?
Those are not just academic queries. They map directly to product experiments and policy choices that will determine whether AI amplifies human judgment or erodes it.
Key takeaways: honest questions and short answers
- Does having AI available make people less likely to say “I don’t know”?
Yes. Across five experiments with 3, 132 participants, access to AI dropped withholding from roughly 36-44% in control groups to single digits in the AI conditions (6% and 3% in Studies 1a and 1b).
- Does confidence change when AI is present?
Greatly. In Study 2 mean confidence increased from 29.6 to 75.9 on a 0-100 scale, even as accuracy fell sharply.
- Do small financial incentives fix the problem?
Partially. The ±$0.10 incentive reduced AI requests and slightly increased abstention, but it did not restore withholding to no‑AI baseline levels.
- Are automatically shown AI answers worse than on‑demand suggestions?
Automatic answers produced nearly the same, or larger, reductions in abstention; Study 4 showed withholding dropped to about 1% when AI content was displayed automatically without incentives.
- Should organizations be worried?
Yes, especially where UI choices and routine access make confident‑sounding AI output the path of least resistance. Design, accountability, and training will determine whether AI strengthens or weakens human judgment.
“Whether human judgment holds up as AI systems spread may depend less on making models more accurate and more on people continuing to recognize and respect the limits of their own knowledge.”, Matthias Bastian, The Decoder (Sep 26, 2026).