AI Medical Scribe Hallucinations: A Working Toolkit for US Physicians
An audit study released this year checked 565 notes from a commercial AI scribe across 142 real patient visits. Two doctors listened to the original audio and compared it against what the AI wrote, line by line. Out of everything they checked, 31.3 percent of notes had a mistake that held up under that kind of scrutiny. That’s roughly one note in three.
This blog walks through what those mistakes look like, where they tend to show up, how to catch them before you sign a note, and what to set up in your practice so you’re not relying on luck. Let’s start with the basics, because not every mistake in a note is the same kind of problem.
Three Kinds of Mistakes, and Why One of Them Is Harder to Catch
| Type | What it looks like | How easy it is to spot |
| Something missing | The AI leaves out a detail the patient mentioned | Easy, you’ll usually notice the gap |
| Something garbled | A name or number comes out wrong | Fairly easy, catch it by reading the line out loud |
| Something made up | The AI adds a finding, diagnosis, or plan that never happened | Hard, it reads like a normal, well-written sentence |
That last one is what people mean by a hallucination. It doesn’t read like an error. It reads like every other sentence in the note, which is why it slips past a quick review. Once you know that’s what you’re up against, the next useful thing is knowing where these mistakes tend to happen, because they’re not spread evenly across the chart.
Where AI Scribes Get It Wrong Most Often
| Area | What tends to go wrong | The evidence |
| Medications | Wrong drug, wrong dose, made-up tapering instructions | An official audit of twenty approved scribes found 12 out of 20 recorded a drug different from what was actually prescribed |
| Phone and video visits | Exam findings written in, even though no exam happened | The same study found this pattern specifically on telephone visits |
| Plans that changed mid-visit | A doctor proposes something, then takes it back, and the note keeps the first version | Documented as its own distinct error type in that research |
| Patient identity fields | Wrong name spelling, wrong birthdate, mismatched details | One of the largest error categories in that same study |
| Mental health topics | A patient brings up sleep, mood, or stress, and it doesn’t make it into the note | The official audit above found 17 out of 20 scribes missed details like this |
Each of these has its own reason behind it, and knowing the reason tells you exactly what to check for. Medications are worth starting with, since they show up as the biggest problem across nearly every study on this.
Why Medications Deserve the Closest Look
Drug names, doses, and refill schedules are packed with specific numbers and terms that sound similar to each other. A model can land close to the right answer without landing on it exactly, and close is often worse than obviously wrong, because a slightly off dose still sounds like something a doctor would write. If a visit involved starting or adjusting a taper, that’s several numbers strung together (starting dose, how often to increase it, the target dose), and each one is a separate chance for something to drift from what was said. The official audit mentioned above is the clearest evidence of how often this specific category goes wrong.
Reading fast costs you the most in this section. Say the line out loud and check it against your memory of the conversation, not against whether it sounds medically reasonable, since a wrong dose usually sounds reasonable too.
There’s a similar pattern with visits that never involved a physical exam at all, and it’s specific enough to deserve its own mention.
The Telehealth Blind Spot
When a scribe writes “lungs clear to auscultation” into a note from a phone call, it isn’t guessing wrong in the usual sense. It’s completing a sentence pattern it learned from thousands of in-person visits, with no sense that this particular call never involved a stethoscope. The audit study found this exact pattern concentrated in telephone consultations specifically. If your practice runs telehealth or phone follow-ups, this isn’t something that happens once in a while. It’s built into how those visits get written up every time, unless someone catches it.
Any exam finding on a remote visit note needs a direct check, the same way you’d double check a lab value you weren’t sure about. Exam findings aren’t the only thing that can show up without happening though. Sometimes it’s an entire plan.
When a Decision Gets Reversed and the Note Doesn’t Notice
Picture a visit where you suggest a referral, then a minute later say you’d rather hold off. The note can still record the referral as if it happened. The same researchers found this as a distinct, repeatable pattern, and it makes sense once you think about how these models process a conversation. They tend to record what was said in the moment, not how the conversation ended up.
This shows up more in visits with a lot of back and forth, weighing a referral, adjusting a treatment plan, talking through options before landing on something final. Pain management, psychiatry, and complex chronic care visits all involve more of this kind of back and forth than a routine same-day visit, so they carry more exposure to this specific mistake.
While you’re checking whether the plan reflects what you decided, it’s worth glancing at something that’s easy to skip entirely.
The Fields Nobody Thinks to Double Check
Wrong identity details, a misspelled name, an incorrect birthdate, a demographic field that doesn’t match anything said, are one of the biggest error categories found in recent research on this. These fields often get skipped during review because clinicians assume they’re pulled automatically from the patient record rather than generated fresh by the same AI writing the rest of the note. They’re not automatic. A quick glance is worth adding to your routine.
Everything so far has been about things the AI adds that shouldn’t be there. Mental health content tends to work the other way around.
Why Mental Health Details Tend to Disappear
This is the one category where the risk runs backwards. Instead of adding something false, the AI tends to drop something true. A patient mentions they haven’t been sleeping, or that getting out of bed feels hard some days. That kind of statement often comes out in plain, everyday language rather than clinical terms, and the model seems less reliable at recognizing it as something worth keeping. The official audit found this specific gap in 17 of the 20 tools it tested.
If a patient brings up mood, sleep, stress, or substance use, check that the note captures it in their own words rather than a generic paraphrase. Here you’re checking for what’s missing, not what’s wrong.
Now that you know where the problems hide, here’s a routine that puts all five of them together into something you can actually run in under two minutes.
A 90-Second Check Before You Sign
| Step | What to do | Time |
| 1 | Read the plan section first, since invented treatments and reversed decisions hide there | 20 sec |
| 2 | Read the medication line out loud, checking the drug, dose, and any tapering instructions | 15 sec |
| 3 | For phone or video visits, check exam findings directly | 15 sec |
| 4 | Glance at the name and birthdate fields | 5 sec |
| 5 | Check that anything the patient said about mood or sleep made it in, in their own words | 15 sec |
| 6 | If a caregiver or interpreter was in the room, confirm statements are attributed to the right person | 10 sec |
Print this and keep it by your screen for the first month. After that it turns into something you do without thinking about it.
A peer-reviewed study published in Annals of Internal Medicine backs up why this kind of close reading matters. Researchers had AI tools and human clinicians write notes for the same five patient visits, then had thirty reviewers score everything blind. AI notes scored lower across every measure tested. For one visit, a case of acute low back pain, the human note scored 43.8 out of 50. The AI note for that same visit scored 20.3. The gap widened even more under background noise, a clinician speaking through a mask, or a patient with a non-native English accent, so give the note extra attention under any of those conditions.
Knowing the order to read in helps. Knowing exactly which sentences to be suspicious of helps even more.
Phrases That Should Make You Slow Down
| Pattern | Example | Why it’s risky |
| Full exam language on a phone or video note | “Lungs clear, abdomen soft, non-tender” | Nobody could have observed that over a call |
| A plan that sounds settled right after a hesitant answer | Patient said “maybe,” note says “agreed to proceed” | Uncertainty often gets smoothed into certainty |
| A referral or test with nothing leading up to it | “Referred to cardiology” appearing out of nowhere | Might be invented rather than discussed |
| A plan you reversed, still shown as active | You said “let’s hold off,” note keeps the original plan | The AI logs what was said, not how it ended |
| A paragraph that feels more polished than the conversation was | Three detailed sentences from a 20-second exchange | Likely filled in rather than transcribed |
| A vague, generic dose where you gave a specific one | “Take twice daily” with no number attached | Sign the exact figure got lost |
Catching mistakes in your own notes is one half of this. Picking a scribe that hands you fewer mistakes to catch in the first place is the other half.
What to Ask Before Choosing a Scribe
A vendor’s advertised accuracy number can tell you less than it seems to. In the government evaluation behind the audit mentioned earlier, accuracy of the notes counted for only 4 percent of the total score used to approve vendors, while a vendor’s physical presence in the region counted for 30 percent. A product can score well in procurement without its notes being especially accurate.
There’s a second reason to be careful with a bare accuracy claim. The same study found that testing identical notes with different review standards produced failure rates ranging from 28 percent to 97 percent. The number a vendor publishes depends heavily on how strict the test was, so two vendors both claiming 99 percent accuracy could have been tested in very different ways.
Score any product you’re considering from 0 to 3 on each line below. Anything under 15 total is worth pushing back on before you sign.
| Question | Score (0-3) |
| Can you hear the exact audio behind any sentence in the note? | ___ |
| Does it handle phone and video visits differently from in-person ones? | ___ |
| Does it flag uncertainty instead of always sounding confident? | ___ |
| Can you see a raw transcript alongside the note? | ___ |
| Do they report error rates by visit type, not one blended number? | ___ |
| Is the patient-facing summary held for your review before it’s sent? | ___ |
| Total | ___ / 18 |
You can also put a product through its own test before signing anything. Record a mock visit with one unclear line, a second speaker playing a caregiver, a specific medication with a taper, a plan you propose and then take back, and a made-up name and birthdate spoken clearly. Run the resulting note through the 90-second check above. Ten minutes of this tells you more than any sales demo will.
While you’re weighing what to ask a vendor, it helps to know what other physicians across the country are asking for too.
What US Physicians Say They Need Before They Trust This
A national survey of 1,692 doctors from the American Medical Association laid out specific conditions physicians want met before they trust an AI tool. It’s a useful list to bring into your own conversations with a vendor.
| What physicians want | How many said it matters |
| Independent testing of safety and accuracy, done on an ongoing basis | 88 percent |
| Not being held personally liable for the AI’s own mistakes | 87 percent |
| Clear assurance about how their data is kept private | 86 percent |
| A real say in whether their practice adopts a given tool | 85 percent |
| A clear explanation of how the tool was built and tested | 80 percent |
If a vendor can’t answer these five points clearly, raise it before you sign, not after. Once you’ve picked a tool and you’re using it day to day, a few practice-wide habits make a bigger difference than people expect.
Habits Worth Building Into Your Day
- Set aside a fixed amount of time to review each note, even a couple of minutes, every time, rather than “whenever there’s time.” That second option tends to disappear the moment your schedule fills up.
- Ask for a raw transcript view if your product offers one, so you can check a claim against the audio instead of trusting a clean summary on its own.
- Update your consent forms to name the AI scribe directly, rather than folding it into general paperwork patients skim past. A simple line works: “This visit may be documented using an AI-assisted scribe tool that listens to and transcribes our conversation. A physician reviews and approves all notes before they’re finalized, and you may decline this at any time.”
- Call your malpractice carrier and ask directly how they handle AI-related claims. Once you sign a note, a mistake in it carries the same weight as one you wrote by hand.
And when you do catch something wrong, here’s exactly what to do about it.
What to Do the Moment You Catch a Mistake
| Situation | Action |
| You catch it before signing | Fix it directly in the note. No further steps needed. |
| You catch it after signing | Log a formal addendum with your name, the date, and exactly what you changed, keeping the original entry legible. |
| It’s a wrong medication dose | Fix it right away, and check whether a prescription already went out based on the wrong number. |
| It already reached the patient through an after-visit summary | Contact the patient directly to correct it, rather than relying on a portal message alone. |
| You keep seeing the same kind of mistake | Report it to the vendor with specific examples, and bring it up at your next team meeting. |
For the addendum, a simple format keeps it clean and audit-proof:
- Date of correction: [date].
- Original note date: [date].
- Corrected by: [your name].
- Original text read: “[incorrect line].”
- Corrected to reflect: “[accurate version].”
- Reason: AI-generated documentation did not match the recorded encounter.
Keep the original wording visible rather than deleting it, since most states want the record to show what was there before and what changed.
Fixing mistakes one at a time only gets you so far though. What actually keeps a practice safe over the long run is a habit that survives past the first few weeks of paying close attention.
The Habit That Fades If You’re Not Careful
Trust in a tool tends to grow faster than the tool improves. You read every note closely for the first few weeks, and by month three you’re skimming, because nothing’s gone wrong yet. That’s usually the moment something slips through.
A few things help this stick. Read each note the way you’d check a colleague’s draft, since that mindset tends to catch more than reading something that already has your name on it. Keep a shared list within your practice of specific mistakes people have caught, so the whole team is watching for the same patterns instead of everyone learning the same lesson alone. Check back on your vendor’s error rate every so often too, instead of assuming the tool you signed up with a year ago still performs the same way, since these products change without always telling you.
Dr. Giriraj Tosh Purohit is an experienced Product Manager and Security officer with a strong background in healthcare technology and management consulting. With expertise spanning clinical workflows, EHR, RCM, Digital Health, and AI-driven products, he has been instrumental in shaping innovative healthcare solutions.