Why Most Ambulatory Practices Are Getting AI Wrong (And What the Data Says Actually Works)
Just a couple of years ago, I watched a small family medicine group spend $180,000 implementing an EHR that their physicians used at about 40% capacity eighteen months later. The vendor had promised transformation. What the practice got was a new way to do the same frustrating documentation, just on a different screen.
I tell that story because AI in ambulatory care is at the same fork in the road. The practices that are getting real, measurable results are not the ones chasing the biggest tools. They are the ones who understood their specific failure points first and matched the right AI layer to each one. This blog is about exactly that.
The Real Problem Is Not Documentation Volume. It Is Documentation Time
Most conversations about AI and physician burnout lead with documentation time. That framing is too broad to be actionable. The more useful question is: which specific documentation tasks are eating time that could be returned to patients or reclaimed as personal time?
The numbers make the problem concrete. The AMA’s 2025 National Physician Comparison Report, drawn from nearly 19,000 responses across 38 states and 106 health systems, found that physician burnout declined for the fourth consecutive year, sitting at 41.9% in 2025, down from 48.2% in 2023. Progress, but unevenly distributed. Emergency medicine hit 49.8% burnout, urological surgery 49.5%, and family medicine 45%. The documentation burden behind those numbers has not eased; the specialties showing the least improvement are the same ones with the heaviest administrative load.
The intent-to-leave number puts a sharper point on it. In 2025, 31.1% of physicians reported a definite, likely, or moderate likelihood of leaving their current organization within the next two years. When physicians leave ambulatory practice for employed positions or early retirement, smaller practices absorb the patient load without the staffing to match.
The question for any practice is not “should we use AI?” It is “which of these specific time sinks do we attack first, and with what?”
Ambient Documentation: What ‘Integration’ Stands For
Physician AI adoption has moved fast. The AMA’s 2026 Physician Survey on Augmented Intelligence, fielded among 1,692 physicians in January and February 2026, found that 81% now use AI in a professional context, more than double the 38% recorded in 2023. A Doximity survey of 3,151 physicians confirmed the pace: daily AI usage jumped from 47% in early 2025 to 63% by January 2026, a 16-point increase in under a year. This reflects broader industry findings. Connext’s 2026 AI Oversight Report revealed that 63% of users report AI outputs are accurate only sometimes or less, with nearly half (46%) noting that fixing errors takes just as much time as doing the work manually.
The adoption numbers look strong, but the outcomes data tells a more complicated story.
The single biggest predictor of whether an ambient AI scribe delivers real time savings is how deeply the tool writes back into the EHR.
There are three tiers of this integration:
Tier 1: The AI transcribes the visit and delivers a draft note. The physician copies and pastes it into the EHR manually. This is the majority of current deployments, and it saves some cognitive effort but very little clock time.
Tier 2: The AI writes the note directly into the EHR but into a generic note field. The physician still has to manually move information to the problem list, medication fields, and follow-up orders.
Tier 3: The AI writes structured, field-level data into the correct EHR locations, including diagnoses, medication changes, and follow-up orders. The physician reviews and confirms. Nothing is copied. Nothing is manually moved.
Only Tier 3 produces the outcomes that show up in peer-reviewed data. A 2025 multicenter study published in JAMA Network Open following 263 clinicians across six health systems found that 30 days of ambient AI scribe use dropped burnout from 51.9% to 38.8%, with significant reductions in after-hours documentation time and cognitive task load.
The most detailed longitudinal picture comes from The Permanente Medical Group, which deployed ambient AI scribes across 7,260 physicians documenting more than 2.5 million patient encounters. A follow-up analysis published in NEJM Catalyst found the system saved physicians an estimated 15,791 hours, with 84% reporting improved patient communication and 82% reporting improved overall work satisfaction.
But not all deployments perform equally. A large April 2026 JAMA study of more than 1,800 clinicians across five academic medical centers found more measured gains, 16 minutes of documentation time and 13 minutes of total EHR time saved per eight-hour care day, with the largest benefits concentrated among the highest-frequency users. The differentiator, consistently, was how deeply clinicians integrated the tool into their workflows.
Before signing any ambient AI contract, ask the vendor one specific question: “Does your platform write structured data to field-level locations in my EHR, or does it write to a general note field?” The answer tells you whether you are buying Tier 3 or Tier 1.
Coding and Billing AI: The Revenue You Are Already Losing
The documentation problem is visible. Physicians feel it every night. The billing problem is quieter but potentially larger in financial impact.
The national average denial rate has climbed to 11.65% in 2026, with high-acuity specialties like orthopedics, oncology, behavioral health, and neurology routinely exceeding 15%. Reworking a single denied claim now costs between $25 and $57 in added overhead.
A 2025 MGMA poll found that 71% of practice leaders now report some AI use in patient visits, yet most RCM operations remain largely manual, meaning the majority of practices are still absorbing denial rework costs through staff labor.
But what our AI-powered autonomous coding can do for that problem is concrete. Our AI Medical Coder reads unstructured physician notes, operative reports, and discharge summaries, then assigns the most specific ICD-10, CPT, and HCPCS codes while flagging denial risks, modifier gaps, and bundling issues before a claim leaves your system. We cut coding time by 50%, score claims against payer-specific denial patterns in real time, and integrate directly with Epic, Cerner, Athena, eClinicalWorks, and other major billing platforms.
The reason the denial reduction is so large is not that AI codes better under pressure. It is that AI codes consistently. Human coders, especially when fatigued or working through a backlog, apply criteria inconsistently. The denial rate for human-coded claims is higher not because coders are bad at their jobs but because volume and complexity make consistency impossible without automation.
For a practice running 20 physicians, even a 20% reduction in coding denials recovers a material share of revenue that is currently being written off or reworked at significant staff cost.
Agentic AI: What It Does That GenAI Does Not
Most AI tools in ambulatory settings today are generative AI tools. They respond when you prompt them. You give them a transcript and they produce a note. You give them a claim and they check modifiers. They require initiation by a human every single time.
Agentic AI is built differently. It monitors, decides, and acts across multiple systems without waiting for a prompt. The distinction is not academic. It changes which problems become solvable.
Here is what a properly configured agentic layer does across a single patient visit in a practice that has deployed it:
Before the patient arrives:
- Pulls the patient’s recent labs, medication list, and reason for visit into a pre-visit brief accessible to the clinician before entering the room
- Checks eligibility against the current insurance record, not the one on file from the last visit
- Stages prior authorization documentation for any procedure or medication the clinician is likely to order, based on the appointment type
During the encounter:
- Ambient AI captures the visit and generates the structured note
- Evidence agents surface real-world clinical data for the specific patient being seen, pulled from patient-level records, directly inside the EHR — without the physician asking for it
After the visit:
- The note is coded with ICD-10 and CPT codes applied
- The claim is submitted
- If a prior auth was triggered, the agent initiates submission, monitors payer status, and escalates to a human only when payer-specific criteria require clinical judgment
Moreover, the AMA’s 2025 Prior Authorization Physician Survey, fielded in December 2025 across 1,000 practicing physicians, found that the average practice completes 40 prior authorizations per physician per week, consuming 13 hours of physician and staff time. 82% of physicians report that patients at least sometimes abandon treatment because of authorization delays, and 94% say prior authorization contributes to burnout. An agentic PA layer does not eliminate that 13 hours entirely, but it eliminates the parts that do not require clinical judgment — which in a typical practice is the large majority of it.
No-Shows: The Revenue Problem Every Practice Underestimates
Patient no-shows cost the US healthcare system an estimated $150 billion annually. For a solo physician practice, that translates to roughly $150,000 in lost revenue per year. The per-appointment cost averages around $200, and no-show rates across US outpatient settings run between 15% and 30% depending on specialty and patient population.
The practices closest to solving this problem are the ones deploying conversational AI that identifies high-risk appointments in advance and acts on them. Healthcare organizations deploying conversational AI for patient engagement are reporting no-show reductions of 25% to 38%, with one health system documenting $804,000 in recovered revenue over seven months from a 28% no-show reduction.
Despite that, only 19% of medical practices currently use AI for patient communication, according to a 2025 MGMA survey, which means adoption remains a real competitive differentiator for the practices that move first.
The mechanism matters. High-risk appointment identification uses historical patterns: appointment lead time, prior no-show history, insurance type, appointment type, and time of day. When a high-risk slot is identified, the system reaches out through the patient’s preferred channel (text, voice, or app), offers easy rescheduling at hours when the front desk is unavailable, and immediately backfills from a waitlist if the patient declines. The front desk only gets involved when the patient requests something outside the defined parameters.
A 2025 Accenture survey found that 79% of patients now prefer digital-first communication with their healthcare providers, rising to 87% among patients aged 25 to 54. The patient preference is there. The question is whether the practice has built the infrastructure to meet it.
The prior authorization burden connects directly here. The same AMA survey shows that 82% of physicians report patients abandoning recommended treatment because of authorization delays. Some of those abandoned treatments show up later as no-shows or late cancellations. Reducing PA cycle time from days to hours does not just improve billing efficiency. It reduces the window in which patients change their minds or lose access.
What Breaks an AI Rollout Before It Starts
I have watched rollouts fail in ambulatory practices that had the right tools, the right budget, and genuine physician buy-in. The failure mode is almost always the same: the practice selected AI to fix a problem but never defined what ‘fixed’ would look like in measurable terms, and so they had no way to know whether it was working or who was responsible when it was not.
The 2026 Doximity Report found that while 94% of physicians are either using or interested in using AI, 71% cite accuracy and reliability of AI outputs as their top concern, and nearly half report that their institution’s AI policies are still evolving or unclear. Physicians who have had no input in tool selection have no ownership of the outcome, and they find workarounds rather than adapting.
The practical implication: the first AI tool a practice deploys sets the adoption trajectory for everything that comes after. A physician who uses an ambient scribe that works exactly as described, requires minimal adjustment to their workflow, and demonstrably reduces their charting time by 6 PM becomes an internal advocate. A physician who uses an ambient scribe that misfires, generates inaccurate notes, or requires more editing than writing from scratch becomes an obstacle to every subsequent AI conversation.
The specific things that need to be defined before any AI tool goes live in a practice:
- Which staff roles are responsible for handling what the AI escalates, and have those people been trained specifically on escalation paths, not just general tool use?
- What specific metric is this tool supposed to move, and by how much, in what time frame?
- Who owns the outcome measurement? (Not the vendor. Someone inside the practice.)
- What is the decision rule for discontinuing if the metric does not move?
The last point is where most rollouts get soft. Practices train staff on how to use the tool. They do not train staff on what to do when the tool hands something back. That gap creates abandoned queues, frustrated clinicians, and a narrative that AI does not work, when the actual problem is that no one defined the human backstop.
The Integration Question That Determines ROI Before You Buy
Every AI vendor you speak with will tell you they integrate with your EHR. The word ‘integrate’ means different things to different vendors, and the gap between those meanings is where most ROI disappears.
There are four specific questions to ask every vendor before a contract is signed:
- Can you show me a live demo on our specific EHR, on a real patient record structure, not a demo environment?
- Does your system write structured, field-level data back to the EHR, or does it write to a free-text note field?
- What happens to data generated by your system if our EHR goes down for two hours? Where does it live, and how is it reconciled?
- What is the FHIR API version you support, and which EHR-specific constraints limit what you can write back?
A vendor who cannot answer the third question in specific technical terms is almost certainly a Tier 1 integration. A vendor who cannot answer the fourth question in a live environment is likely to underperform post-deployment.
The practices that have achieved the outcomes in the peer-reviewed literature did so with tools that solved the integration problem before deployment, not after. The research showing meaningful daily documentation savings per clinician assumes Tier 3 integration. Most vendor demos assume the same. What arrives post-contract is often Tier 1.
Final Thoughts: The Staff Adoption Problem Is a Procurement Problem
The Doximity data shows AI adoption rising to 63% of physicians by January 2026, but it also shows something more nuanced: family medicine physicians, once they adopt, use AI daily at a rate of 88%. The challenge is not sustained use. It is the on-ramp.
This means the stakes on the first tool selection are higher than they might appear. Choosing the safest-looking vendor or the cheapest option is not conservative. It is high-risk for adoption velocity across the whole practice.
The practices with the highest AI utilization rates share a common pattern: they piloted a single tool with a small group of volunteers, measured the outcome honestly, made the tool available broadly only after those volunteers became advocates, and tied the rollout to a specific, visible problem that every physician in the practice already agreed was painful.
So, the question is whether your practice selects a tool in a way that gives it the best possible chance of becoming the first kind of story instead of the second.

Dr. Giriraj Tosh Purohit is an experienced Product Manager and Security officer with a strong background in healthcare technology and management consulting. With expertise spanning clinical workflows, EHR, RCM, Digital Health, and AI-driven products, he has been instrumental in shaping innovative healthcare solutions.