
In a stark illustration of the risks posed by artificial intelligence in clinical settings, a patient in England was falsely informed that they had demyelination, the nerve damage associated with conditions such as multiple sclerosis. The actual test result had read "null demyelination" — meaning no demyelination was detected — but an AI scribe had dropped the word that reversed the entire meaning. This incident is one of several documented in a new warning from Healthwatch England, the statutory patient watchdog.
The finding has sent ripples through the National Health Service, where AI-powered scribes have been rapidly adopted as a solution to the growing burden of clinical documentation. These tools sit in the consulting room, listen to the conversation between clinician and patient, and automatically generate the notes that go into the permanent medical record as well as the letters that are sent to patients. According to Healthwatch England, twenty-seven different AI scribe products are currently in use across the health service in England.
When Fluent Notes Are Wrong
The errors collected by the watchdog are mundane in form but serious in effect. In one case, a scribe swapped a prescribed drug for a different medication with a similar name — the kind of confusion that pharmacology spends considerable effort trying to design out of the system. In another, the AI summary omitted a consultant's instruction that the patient should seek a repeat prescription for migraine medication. A third summary recorded a doctor telling a patient to continue taking Prozac, when that doctor had neither prescribed it nor even discussed it during the consultation.
What links these failures is that nothing looked broken at first glance. A fluent, plausible note is exactly what these systems are designed to produce, which is why an error can survive the quick glance a busy clinician gives it before signing off. The watchdogs warn that inaccuracies may persist in patients' records indefinitely if the patient does not catch them. This places the last line of defence on the person least equipped to know what the note should have said — the patient, who may not be familiar with medical terminology, drug names, or the full clinical context.
Concerns from Experts and Advocates
Rachel Power, chief executive of the Patients Association, is among those raising concerns, alongside clinicians such as London GP Shier Ziser Dawood and Charlotte Blease, a researcher at Uppsala University in Sweden. The objection is not to the technology itself but to its arrival without a safety net. There is genuine potential for AI scribes to reduce the administrative burden on doctors and improve the patient experience by allowing clinicians to focus on the conversation rather than on typing notes. But that potential is undermined when errors go undetected and unrectified, and when there is no systematic process for checking the accuracy of AI-generated documentation.
The promise of ambient documentation is significant. Doctors across the world spend hours each day entering data into electronic health records, often sacrificing time that could be spent with patients. AI scribes offer the possibility of a consultation where the doctor is fully present, taking a history, examining, and explaining, while the technology quietly captures the details. This is an attractive vision, and it explains why so many clinicians have been eager to experiment with these tools. However, the rush to adopt has outpaced the development of robust safeguards.
The Regulatory Gap
A central problem is the absence of England-wide oversight for these tools. The Medicines and Healthcare products Regulatory Agency (MHRA) has not classified AI scribes as medical devices, which leaves them outside the regulatory regime that would test them for safety and effectiveness before deployment. This means that, unlike pharmaceuticals or traditional medical devices, AI scribes are not subject to mandatory pre-market scrutiny.
The MHRA did publish guidance in August clarifying where the line sits. According to that guidance, a system that only transcribes what was said is not a medical device, whereas a system that suggests a diagnosis or a treatment may well be. This puts a great deal of weight on how each product is described by its vendor. If a scribe is marketed purely as a passive transcriber, it avoids the regulatory process that a scribe marketed as a clinical assistant would have to complete. The incentive that follows is obvious: vendors may choose to downplay the decision-support capabilities of their products to steer clear of stringent oversight.
This classification gap is not merely a technicality. It has tangible consequences for patient safety. Without the requirement for clinical validation, there is no standardised way to measure the accuracy of these tools or to compare them against each other. Clinicians and patients alike are left in the dark about the reliability of a system that is quietly becoming a permanent fixture in the consulting room.
Different Failure Modes
Transcription error is also a different failure from the one most AI safety work anticipates. In the cases highlighted by Healthwatch England, nobody was misled by a hallucinated fact — a fabricated symptom or a made-up diagnosis. Instead, a real sentence was rendered slightly wrong, and slightly wrong is sufficient when the sentence names a drug or a clinical finding. The demyelination case is a telling example: the original test result was accurate, but the AI's omission of the prefix "null" turned a negative result into a positive one, with potentially devastating psychological and clinical consequences.
This discrepancy highlights a broader challenge in the field of artificial intelligence in healthcare. While much attention is paid to the dramatic failures of AI systems — such as generating false information with confidence — the quieter, more insidious errors of subtle mis-transcription may be more common and equally dangerous. In a high-stakes environment like healthcare, the difference between a correct and an incorrect note can be the difference between proper treatment and medical harm.
Another complication is that language models are not designed to be verbatim recorders. They are built to generate probable text, which means they often fill in or simplify gaps in a way that a human transcriber would not. This is especially concerning in medicine, where a single word, such as "not" or "null," can completely alter the clinical picture. The risk extends beyond drug names and diagnoses to dosages, allergies, and follow-up instructions.
Why AI Scribes Spread So Quickly
None of which addresses why these tools spread so quickly across the NHS. Clinical documentation is the administrative burden that doctors complain about most. Studies have shown that clinicians spend a significant portion of their day — sometimes hours — updating electronic health records, typing notes, and managing paperwork. A system that reliably removes an hour of typing a day will be adopted whether or not anyone has formally assessed it. The pressure on general practitioners and hospital doctors is so intense that any tool promising to alleviate the workload is likely to be welcomed with open arms.
The problem is what happens in between those two facts. A tool adopted for its speed, unassessed because of how it is categorised, producing a document that becomes the permanent clinical record — that is a chain in which no single link is obviously anyone's responsibility. The clinician who signs off the note may not have time to verify every detail. The vendor may have designed the product to optimise for fluency rather than verbatim accuracy. The regulator may not have the authority to intervene because of the device classification gap. And the patient, who is most motivated to spot errors, is often not shown the note until it arrives in the post, and may not understand the medical jargon well enough to catch subtle mistakes.
A Deliberate Safety Mechanism
The remedy that the watchdog points to is unglamorous and probably right. Patients should be told when a scribe is being used, and they should be given their notes to check. This would turn an accidental safety mechanism into a deliberate one. Instead of relying on the unlikely chance that a patient happens to read their letter carefully and notices an error, the NHS should build a formal process in which patient review is an expected step. This could be as simple as providing a printed copy of the AI-generated notes at the end of the consultation, or sending a digital copy with a clear invitation to flag any inaccuracies.
Some healthcare providers have already started to implement such practices, but it is far from universal. The variation in how AI scribes are deployed across different trusts and practices is itself a cause for concern, as it means that patient safety depends on where they happen to be treated. A national standard for the use of AI scribes, including mandatory patient notification and a clear mechanism for correcting errors, would go a long way toward addressing the current gaps.
There is also an argument for stronger regulatory oversight. If AI scribes are capable of influencing clinical decisions, they should be subject to the same scrutiny as other clinical decision-support tools. The MHRA's August guidance is a start, but it may be insufficient to capture the full range of risks. As these systems become more sophisticated, integrating not just transcription but also summarisation, coding, and possibly even diagnostic suggestions, the line between a passive tool and an active medical device will become increasingly blurred.
The experiences of other countries may offer lessons. In the United States, several health systems have deployed ambient clinical documentation tools with varying degrees of success, and some have implemented mandatory audit trails and clinician training programs. In Scandinavia, where digital health records are long-established, there is a strong culture of patient access to records, which has helped to surface errors. The NHS could look to these examples as it considers how to regulate and deploy AI scribes safely.
The Current Arrangement
Twenty-seven products, no device classification, and a check performed by whoever happens to read their letter carefully: that is the current arrangement in the NHS in England. It is an arrangement that puts an extraordinary burden on patients and relies on the goodwill of individual clinicians to catch errors that the AI system has introduced. The warning from Healthwatch England is a reminder that technological innovation in healthcare must be accompanied by appropriate safeguards. The promise of AI scribes is real, but so are the risks. Until a robust regulatory framework is in place, and until patients are systematically involved in checking the accuracy of their records, these tools will continue to operate in a grey zone where a wrong drug name or a missing "null" can have life-altering consequences.
