When AI scribes write the clinical story: errors, oversight gaps, and what leaders must do now
One patient opened their hospital record and saw the word “demyelination” attached to their name. It sounded like a life-changing diagnosis. The hospital later amended the entry to “null demyelination, ” but only after the patient, an NHS health professional, queried the record and asked not to be named.
That case sits at the heart of a warning from Healthwatch England: AI “scribe” tools that listen to clinician-patient consultations and generate summaries are producing clinically significant errors, wrong diagnoses, wrong drugs, and omissions, and those errors are appearing in patient records. Healthwatch says it has heard “multiple stories from patients who have noticed these errors when a health professional hasn’t.”
How these failures show up
Healthwatch collected several concrete failures that illustrate the danger:
- An AI scribe recorded “demyelination” when the correct MRI interpretation indicated there was none; the hospital later corrected the record to read “null demyelination.”
- An AI tool confused one prescribed drug for another with a similar name, a patient, not the clinician, spotted the mismatch.
- An AI-generated summary letter omitted that the consultant had asked the patient to seek a repeat prescription for migraines from their GP.
- Clinicians have reported “hallucinations, ” where a system invents facts or instructions not present in the consultation, for example recording that a doctor told a patient to “continue their Prozac” when that drug was never prescribed or discussed.
These are not harmless typos. Misstating a diagnosis, misstating medication, or omitting treatment instructions can alter care pathways, prompt unnecessary investigations, or seriously distress patients.
Where the risks cluster
From the cases and clinician surveys, high-risk scenarios are already visible:
- Complex consultations with multiple problems or multiple speakers. That creates more opportunity for misattribution and omission.
- Consultations involving non-native English speakers or strong regional accents. Speech-recognition errors rise in predictable ways; local reports include complaints that an AI receptionist “did not understand their strong Yorkshire accents.”
- Medication and diagnosis names that sound alike. Small transcription errors become clinically significant.
- Situations where clinicians rely on the scribe and do not thoroughly review or correct the output before it is finalised into the record.
What clinicians and researchers are reporting
Adoption is spreading. Clinicians across England are using a range of AI scribe products. Experience is mixed. London GP Dr Shier Ziser Dawood warns these systems reduce clerical load but “may be a double-edged sword, ” recounting examples of hallucinations and cautioning that replacing note-taking could create expectations for clinicians to see more patients, even though UK GP consultation times are already short.
Academic researchers are blunt. Dr Charlotte Blease of Uppsala University told reporters: “AI can and does make mistakes.” Her survey of 1, 003 UK GPs found that more than half believed ambient AI records were more accurate than their own notes, even as clinicians reported error patterns were more likely when consultations included multiple participants, when patients had complex medical histories, or when English was not the patient’s first language.
A regulatory and governance gap
Part of the alarm is governance. The Medicines and Healthcare products Regulatory Agency (MHRA) has taken a nuanced stance: its guidance indicates some transcription or administrative tools may fall outside the definition of a medical device, while software that informs diagnosis or treatment decisions is more likely to be classed as a device and subject to regulation. That distinction matters because where a tool sits determines whether it faces England-wide device-style oversight.
That regulatory nuance sits uneasily next to government ambitions. The UK government’s 10-year health plan expects AI scribes to “liberate staff from their current burden of bureaucracy and administration, freeing up time to care and to focus on the patient.” Healthwatch and patient groups say the present rollout exposes gaps in reporting, correction pathways, and vendor transparency that must be addressed if that promise is to be realised safely.
“Trust and confidence in this technology depend on good communication and genuine partnership with patients and right now both are missing.”, Rachel Power, Patients Association chief executive
“Healthcare has never been error-free. But our findings show the urgent need for clarity over how patients can report and get corrected any mistakes made by AI scribing tools or the professionals that use them.”, Healthwatch spokesperson
Practical controls for healthcare leaders
Health systems have to balance productivity gains against safety and liability. The sensible path is cautious deployment with tight human-in-the-loop controls and transparent accountability. Practical, measurable steps clinical leaders can implement immediately:
- Require clinician sign-off before notes are finalised. The clinician who saw the patient should review and sign (initial or electronic signature) any AI-generated note before it becomes a permanent record. A target: 100% clinician review prior to finalisation, with clear audit trails.
- Establish a rapid patient correction pathway. Notify patients when an AI-assisted note is used; provide a one-click “flag this item” button in patient portals or by phone; set a service-level target, for example clinical review of flagged items within 48 hours, and an auditable feedback loop to the vendor.
- Audit and publish error metrics. Track metrics such as transcription error rate per 1, 000 words, percent of records with clinically significant errors, time-to-correction after detection, and error distribution by consultation type, speaker, and language. Share summaries internally at least weekly during early deployment.
- Demand vendor transparency and SLAs. Require vendors to provide product versioning, transcription confidence scores, error logs, incident timestamps, and a clear escalation path. Include contractual obligations for timely patching, incident response, and data governance.
- Train clinicians and staff. Build short modules on common AI failure modes (hallucinations, misattribution, homophones) and practical spot-checking techniques. Reinforce ownership: the clinician remains legally and ethically accountable for the record.
- Start small and measure impact. Use pilot deployments in controlled settings, with explicit KPIs for safety and productivity, for example minutes saved per consultation versus number of corrections needed. Scale only when error rates are acceptable and governance is strong.
These measures won’t eliminate risk, but they shift the system from passive adoption to active governance. They create an environment where the technology’s administrative benefits can be realised while limiting the chance of harm, and they give patients a clear route to correct their records when something goes wrong.
Key questions executives should be able to answer
-
Can AI scribes make clinically significant mistakes?
Yes. Healthwatch England has compiled multiple patient-reported cases where AI scribes produced wrong diagnoses, confused drug names, or omitted critical instructions.
-
Are clinicians always catching these errors before they affect care?
No. Healthwatch reports several examples where patients, not clinicians, identified the mistakes.
-
Is there national, device-style regulatory oversight for AI scribes in England?
Not uniformly. MHRA guidance suggests some administrative transcription tools may fall outside medical-device rules, while software that influences diagnosis or treatment is more likely to be regulated, leaving a governance gap for many scribe deployments.
-
Do clinicians trust ambient AI records?
Trust is mixed. A survey of 1, 003 UK GPs conducted by Dr Charlotte Blease found that more than half believed ambient AI records were more accurate than their own notes, even though errors and ‘hallucinations’ are a known failure mode.
-
What are the first practical steps organisations should take?
First 30 days: require clinician sign-off for all AI notes and publish weekly internal error summaries. Within 90 days: secure vendor transparency agreements (error logs, SLAs), implement a patient correction pathway with a defined review target, and run focused audits on high-risk consultation types.
AI scribes can reduce the clerical burden on clinicians, but they introduce new, less-predictable error modes that demand new governance. The right posture for leaders is not reflexive embrace or blanket rejection. It is careful, measured adoption with clear human accountability, strong reporting, and mechanisms that let patients correct the record quickly when machines get it wrong.