Autonomous clinical AI and the human final sign-off: evidence, risks, and a safer policy path

On Aug 18, 2026, an opinion in JAMA by bioethicist Ezekiel Emanuel and Neal Khosla (CEO of Curai Health) argued that autonomous clinical AI, meaning systems that make cognitive medical decisions without a mandatory human final sign‑off, may soon outperform doctor‑plus‑AI teams on key reasoning tasks. Their prescription is provocative: regulators should avoid locking a human “final sign‑off” into law when the machine is demonstrably superior.

Readers should weigh that prescription alongside the authors’ affiliations and potential incentives: Neal Khosla leads Curai Health, and the piece notes investor ties (Vinod Khosla is an investor in both OpenAI and Curai). Those links matter when the paper’s prescriptions align with commercial opportunities.

What the JAMA authors point to

The JAMA piece pulls together several recent findings the authors say support selective autonomy on cognitive tasks. Three patterns they highlight:

  • Several retrospective and simulated comparisons since 2024 report that large language models and diagnostic orchestration systems can match or exceed clinicians on narrowly defined reasoning tasks (taking history, diagnosing, choosing tests, treating to guidelines, and managing chronic disease, the five core reasoning tasks the authors name).
  • Some experiments suggest human reviewers can degrade performance when the AI is already superior, a pattern consistent with automation bias and deskilling (the authors cite a meta‑analysis of 106 experiments in support of this point; the meta‑analysis should be reviewed directly for scope and clinical relevance).
  • Evidence that clinician skills can atrophy with routine reliance on AI tools, the authors cite, for example, a Lancet paper on colonoscopy outcomes as an illustration that procedural or cognitive skill erosion can occur when clinicians rely on assistive systems.

Specific items mentioned in the JAMA piece (summary and citation status):

  • Google’s conversational system AMIE, reported to have “scored higher than primary care doctors in almost every category during simulated patient conversations.” (As reported by the JAMA authors; a primary, peer‑reviewed evaluation was not included in the materials I reviewed and should be obtained for independent appraisal.)
  • A head‑to‑head claim that “ChatGPT o3” across 377 complex cases named the correct diagnosis first 60% of the time vs 15.9% for 20 internists. (Reported in the JAMA piece; the underlying study and methods need to be inspected for information parity, case selection, and whether the model had seen similar data in training.)
  • A Microsoft “diagnostic orchestrator” that reportedly found the correct diagnosis under budget constraints about four times as often as doctors, and at lower cost. (Reported outcome; seek Microsoft’s evaluation report for dataset and cost assumptions.)
  • A study cited where GPT‑4 alone scored 92% on diagnostic reasoning while doctors with access to the same model scored 76%. (Reported in the opinion; obtain the primary study to verify scoring rubric, case mix, and study design.)
  • At least one peer‑reviewed, real‑patient retrospective study supports stronger LLM performance in constrained settings: a German emergency‑department study comparing GPT‑4 to resident physicians (LMU Munich) found GPT‑4 outperformed residents on a 100‑case retrospective chart comparison (see: “Evaluating ChatGPT’s Diagnostic Accuracy in Emergency Medicine, ” PMC link: https://pmc.ncbi.nlm.nih.gov/articles/PMC11263899/). This study increases ecological validity but remains single‑center and retrospective.

Strengths, why the argument matters

  • Empirical traction: Multiple studies now show LLMs and orchestration systems can reach or exceed clinician performance on limited tasks. That changes the baseline for regulatory thinking.
  • Human factors evidence: Automation bias and deskilling are real. If clinicians routinely overrule superior model outputs because of process mandates or poor interface design, patients could be harmed.
  • Policy urgency: Locking a rigid human‑final‑signoff requirement into regulation risks ossifying workflows that may become demonstrably inferior. Regulators should design standards that can adapt to evidence instead of freezing a specific process model.

Limitations and open questions

  • Evidence quality and generalizability: Many of the boldest results come from simulations, vignettes, or retrospective datasets. Those designs are prone to selection bias, information parity problems (models sometimes get cleaner or more complete inputs than clinicians), and potential data leakage. Prospective, multicenter randomized evaluations are scarce.
  • Narrow metrics: Studies often measure diagnostic accuracy in isolation (top‑1/top‑3 accuracy) rather than whole‑patient outcomes (mortality, avoidable harm, care coordination, patient experience).
  • Handoff fragility: Translating a high‑performing model into practice depends on interfaces, uncertainty communication, escalation rules, and clinician workflows. Poorly designed handoffs, such as ambiguous confidence signals or unclear escalation thresholds, can turn a safety net into a new failure mode.
  • Distinct autonomous failure modes: Autonomous cognitive systems introduce problems different from human error, including misleading but plausible outputs (“hallucinations”), dependence on connectivity, model drift, and security vulnerabilities. These require engineering and regulatory mitigations that differ from those for human error.
  • Ethical, legal, and trust issues: Liability allocation, informed consent for AI‑led care, and patient trust remain unresolved and will shape adoption far more than raw accuracy numbers.

“When the machine is clearly ahead, the doctor doing the checking turns from a safety net into a source of error.”, Ezekiel Emanuel et al., JAMA opinion (Aug 18, 2026)

“AI‑only care is the ‘economy class’ of medicine.”, Robert Wachter (critical perspective cited by commentators)

What this means for health systems, payers, and regulators

Leaders need a middle path that neither reflexively preserves human final‑signoff nor rushes to autonomous deployment based on vignettes. The practical program looks like three parallel tracks: (1) generate prospective evidence where autonomy is proposed, (2) build the engineering and monitoring infrastructure that makes autonomy safe in production, and (3) update legal and payment systems so incentives and liabilities align.

  • Require prospective, pre‑specified evaluations. Mandate prospective trials or robust real‑world evaluations (stepped‑wedge, cluster RCTs, or phased rollouts) with predefined endpoints and subgroup analyses before authorizing autonomous operation in routine care.
  • Insist on detailed reporting from vendors. Require vendor disclosure of training data provenance, model versioning, calibration metrics, known failure modes, and independent performance audits across subgroups (race, age, sex, comorbidity).
  • Build monitoring and rollback mechanisms. Deploy real‑time dashboards, canary testing, percentage‑shadow deployments, and automatic rollback triggers tied to prespecified performance or safety thresholds.
  • Update contracts and liability frameworks. Create contracts that allocate responsibility among vendor, deployer (health system), and clinicians, and engage payers to align reimbursement with validated outcomes.
  • Redesign clinical roles and training. Train clinicians for system oversight, exception management, and complex longitudinal care, and preserve procedural volumes and cognitive practice where needed to avoid deskilling.

Operational metrics leaders should demand

  • Primary clinical endpoints (not just top‑1 accuracy): patient outcomes, avoidable adverse events, time‑to‑diagnosis, and downstream utilization.
  • Calibration and uncertainty metrics: expected calibration error (ECE) and how confidence scores correlate with true accuracy across clinical subgroups.
  • Subgroup performance: disaggregated accuracy and harm rates by race, age, sex, language, and comorbidity, with prespecified non‑inferiority or superiority margins set by clinicians and ethicists.
  • Safety triggers and rollback thresholds: maximum allowable drift in key metrics, unacceptable subgroup delta, and latency to rollback.
  • Operational KPIs: percent of cases routed for human review, median clinician override rate with rationale logging, and incident‑reporting latency.

Checklist before you remove clinicians from a workflow

  • Is there a prospective, independent evaluation (RCT or robust rollout) demonstrating non‑inferior or superior patient outcomes in the target population?
  • Are confidence estimates calibrated and actionable, i.e., do they reliably identify cases for human review?
  • Is there continuous monitoring with automated canaries and clear rollback triggers?
  • Are subgroup analyses acceptable to clinical and equity stakeholders, with mitigation plans for detected performance gaps?
  • Have contracts and malpractice frameworks been updated to allocate liability and ensure victims of harm can obtain redress?
  • Is there a reskilling plan for clinicians whose everyday cognitive tasks will shift to AI, including preserved procedural volumes where competence is critical?

Key takeaways, questions a curious leader should ask (and short answers)

  • Can AI already outperform clinicians on core diagnostic tasks?

    Some studies, including retrospective and simulated comparisons and at least one peer‑reviewed retrospective ED study (LMU Munich; see PMC link above), show LLMs can outperform clinicians on narrow diagnostic tasks. Many of the most dramatic claims cited in the JAMA opinion come from retrospective or simulated work and require independent verification of the primary sources.

  • Does that mean humans should be removed from clinical decision‑making now?

    The JAMA authors argue regulators should avoid mandating a human final sign‑off when the machine is clearly superior, but current evidence is heterogeneous and often limited to non‑prospective designs. Removing humans should be conditional on strong, prospective evidence and robust production safeguards.

  • What are the main risks of letting autonomous AI operate without mandatory human sign‑off?

    Key risks include misleading but plausible outputs (“hallucinations”), distributional drift, security vulnerabilities, degraded clinician skills over time, and unequal performance across populations, each requiring specific mitigation and monitoring.

  • Who disagrees with the JAMA recommendation?

    Physician organizations such as the AMA and the American College of Physicians, and commentators like Robert Wachter, favor AI as augmentation and caution about replacing clinician judgment wholesale, concerns rooted in safety, ethics, and trust.

  • What should regulators and health systems do next?

    Prioritize prospective evidence requirements, mandate vendor transparency, require continuous monitoring and rollback abilities, update liability and payment frameworks, and fund clinician reskilling tied to measurable milestones.

Three priority actions for executives

  • Fund prospective pilots with clear, clinically meaningful endpoints and mandatory subgroup analyses before scaling autonomy.
  • Require vendor transparency on training data, calibration, and independent audits; implement shadow deployments and canary releases with automated rollback triggers.
  • Negotiate contracts that allocate liability clearly and partner with payers to align reimbursement to verified clinical outcomes, not simply cost savings on paper.

Three red flags that should stop a rollout immediately

  • Unexplained model drift or sudden performance decline in production dashboards.
  • Significant performance gaps for any clinically relevant subgroup without a mitigation plan.
  • Opaque vendor claims (no independent audit, unavailable primary study data, or undisclosed training provenance).

Policy and engineering will both determine whether autonomy becomes a path to better care or a brittle shortcut. The JAMA authors raise an important counterpoint to reflexive calls for human final‑signoff. Their argument should force regulators and leaders to focus on outcomes, evidence standards, monitoring, and incentives now, not after products are entrenched. Do not treat autonomy as an on/off switch; treat it as a high‑stakes experiment that requires transparent evidence and industrial‑grade safety engineering before it becomes standard practice.