The 69/22 Gap: Why Healthcare Runs on Generative AI but Stalls on AI Agents
Technology Blogs

The 69/22 Gap: Why Healthcare Runs on Generative AI but Stalls on AI Agents

Amreetanshu Kumar Sinha
Trainee Software Engineer

I. The Number That Should Be Getting More Attention

A. What the Data Actually Says

There’s a statistic from NVIDIA’s 2026 State of AI in Healthcare and Life Sciences report that deserves more scrutiny than it’s getting: roughly 69% of healthcare organizations report using generative AI, while only about 22% report using AI agents.

That’s a three-to-one gap between organizations that have adopted AI and organizations that have adopted AI that does anything.

It’s tempting to read this as a maturity curve, agents are newer, so adoption lags, and the number will catch up on its own. That reading is comfortable and mostly wrong. Generative AI went from novelty to standard clinical infrastructure in about two years. Ambient documentation was a pilot at a handful of academic medical centers and is now the most widely deployed AI category in healthcare by a significant margin. Healthcare is clearly capable of adopting AI quickly when the path is clear.

So the question isn’t why agents are slow. It’s what specifically is in the way.

B. Where the Adoption Actually Stalls

Industry analysis through 2026 points consistently at the same set of causes. Agentic adoption remains limited because action-based use cases require deeper system integration, higher access rights, and stricter governance than advisory tools do, which makes them substantially harder to deploy at scale.

Read that list again, because none of those items are model problems:

  • Deeper system integration
  • Higher access rights
  • Stricter governance

These are platform engineering problems. They’re the problems that show up after the demo works, when someone in security asks what happens if the agent is wrong at 2 a.m. on a Saturday. For teams already running AI agents in production, none of this will be a surprise, for everyone else, it’s the whole project.

Generative AI vs. agent adoption in healthcare organizations, 2026
Figure 1 Generative AI vs. agent adoption in healthcare organizations, 2026

II. What Actually Changes When AI Acts Instead of Advises

A. The Read/Write Boundary

Most healthcare AI in production today is read-only. An ambient scribe listens and produces a draft. A risk model scores a patient and surfaces a number. A summarizer condenses a chart. In every case, a human takes the output and decides what to do with it.

The system of record never changes unless a person changes it.

An agent breaks that pattern. A scheduling agent moves an appointment. A prior-auth agent submits a packet. A care-gap agent sends a notification to a patient. The write happens without a human initiating that specific action, which is precisely what makes agents valuable, and precisely what makes them hard.

Once you cross the read/write boundary, a whole set of questions that never applied to your summarizer suddenly do:

  • Attribution: Who performed this action, in the audit log, when the actor was software?
  • Authorization: Under whose clinical authority did it happen?
  • Reversibility: Can this be undone, and does undoing it leave a clean trail?
  • Blast radius: If the agent’s reasoning is wrong, does it affect one patient or a panel of four thousand?

B. The Regulatory Line Nobody Wants to Cross First

The regulatory picture reinforces the same boundary. As of April 2026, the FDA had cleared zero agentic clinical AI systems. ARPA-H’s ADVOCATE program is working to establish the first authorization pathway through a multi-year process for a cardiovascular care agent.

Meanwhile, ambient clinical documentation has scaled across health systems, largely because it documents rather than diagnoses, which usually keeps it outside medical device regulation.

That contrast explains a lot of the 69/22 gap on its own. The advisory tools found a path to deployment that didn’t require anyone to be first through an unmapped regulatory door. Agentic clinical tools haven’t yet. And most organizations, reasonably, would rather not fund the mapping expedition.

The practical takeaway for product teams: the line that matters isn’t AI or no AI. It’s advise or act. Where your feature sits relative to that line determines your compliance surface, your liability exposure, and your time to market far more than your model choice does.

Advisory AI vs. agentic AI, regulatory and integration surface compared
Figure 2 Advisory AI vs. agentic AI, regulatory and integration surface compared

III. The Five Things That Have to Exist Before an Agent Ships

This is the part that rarely makes it into AI strategy decks, because it isn’t strategy. It’s plumbing. But it’s where agent projects actually die.

A. Identity: Your Agent Doesn’t Have a User

Every healthcare system’s access model assumes a human principal. A clinician logs in, their role determines their rights, and their actions are attributed to their identity. That model has decades of regulatory and operational assumptions built on top of it.

An agent doesn’t fit. It isn’t a user, and it isn’t quite a system integration either, it acts on behalf of a clinician, sometimes, for some actions, within limits.

The lazy solution is a service account with broad rights. It works in the pilot and fails the first security review, because now every action the agent takes is attributed to a generic identity with more permissions than any single human in the organization.

What you actually need is a delegation model: the agent operates under a scoped grant from a specific clinical authority, that grant is time-bound and revocable, and the audit trail records both the agent and the human whose authority it acted under. SMART on FHIR’s scoping model gives you a reasonable starting vocabulary here, but it was designed for apps a user launches, extending it to autonomous, long-running agents takes deliberate design.

B. Permissions Scoped to the Resource, Not the System

“The agent has FHIR access” is not a permission model. It’s the absence of one.

A medication adherence agent needs to read MedicationRequest and write Communication. It has no business reading psychiatric notes, and it should be structurally incapable of doing so, not merely instructed not to.

Concretely, this means:

  • Per-agent scopes, not per-application: Two agents in the same product get different grants.
  • Resource-type and often resource-instance granularity: Read access to this patient’s observations, not the population’s.
  • Separate read and write scopes, always: Most agents need far more read than write.
  • Deny by default: New capability requires a new grant and a new review.

This is ordinary least-privilege thinking. It just gets skipped constantly, because during development it’s friction and the agent works fine without it.

C. Audit Trails That Reconstruct Reasoning, Not Just Results

Standard audit logging answers what happened. For agents you need to answer why, and you need to answer it months later, potentially to a regulator or an attorney.

That means capturing, for each agent action:

  • The inputs the agent had access to at decision time
  • The specific model and version that produced the decision
  • The intermediate steps, if the agent chained multiple calls
  • The policy or guardrail evaluations that permitted the action
  • The human approval, where one was required, including who and when

This is meaningfully more data than most logging pipelines are built for, and it needs the same retention and access controls as PHI, because much of it is PHI. Retrofitting this later is painful. Build it into the first agent.

D. Reversibility and Containment

Before an agent goes live, someone should be able to answer two questions in concrete terms:

  1. If this action is wrong, what’s the procedure to reverse it?
  2. If the agent is systematically wrong, how many records does it touch before anyone notices?

Question two is the one that gets skipped. A rate limit is a safety control. A daily cap on autonomous actions is a safety control. A staged rollout that starts at one clinic is a safety control. These aren’t performance optimizations, they’re what keeps a bad deployment from becoming a reportable event.

E. Detection for Errors of Omission

This is the least intuitive requirement, and possibly the most important.

Most AI quality assurance is built to catch the model saying something wrong, hallucinations, fabricated citations, incorrect codes. But a 2025 evaluation of clinical agent behavior found that over 80% of severe errors were failures to flag danger, rather than incorrect statements. The system’s mistake was silence.

Almost no standard QA pipeline detects silence. You can’t diff a missing alert against an expected output if your test set only contains cases where something was said.

Building for this means constructing evaluation sets specifically around cases where the correct behavior is escalation, the patient who should have been flagged, the interaction that should have blocked, the result that should have paged someone. Then you measure how often the agent stayed quiet when it shouldn’t have.

There’s a related human factor worth designing against: automation bias. When an agent produces a fluent, well-structured draft, reviewers approve it faster and scrutinize it less. A human-in-the-loop step that becomes a reflexive click is not a safety control, it’s a liability transfer with extra steps. If your workflow depends on genuine review, the interface has to make review genuinely necessary.

Five infrastructure prerequisites for agent deployment in healthcare
Figure 3 Five infrastructure prerequisites for agent deployment in healthcare

Is your platform's data and permission layer ready to support autonomous workflows?

IV. Where Agents Are Actually Working Right Now

The organizations closing the 69/22 gap aren’t doing it with ambitious clinical autonomy. They’re doing it in administrative workflows where the governance requirements are manageable and errors are easy to detect and contain.

Related read: AI Agents in Healthcare: Top Use Cases & Leading Solutions

A. The Categories With Real Traction

1. Prior Authorization

The strongest current example, and notably unglamorous. Prior auth is repetitive and rules-driven: gather data from multiple sources, validate it, assemble supporting documentation, submit. The output is checkable, the failure mode is a rejection rather than a clinical harm, and the volume is high enough that automation pays back quickly.

2. Eligibility and Benefits Verification

Similar profile. Deterministic rules, external systems that need to be queried and reconciled, and an output that’s verifiable before it affects anything clinical. This is the workflow behind InsureVerify AI.

3. Documentation Drafting and Coding Support

Sits right at the advisory boundary. The agent assembles and structures, a human signs. The signature is the control, and it’s a real one as long as the review is real.

4. Structured Patient Outreach

Post-discharge check-ins, remote patient monitoring data collection, adherence reminders. The agent’s autonomy is bounded to sending and collecting; anything clinical routes to a human. The critical design question is escalation: when a patient response indicates deterioration, what happens, how fast, and to whom?

B. The Pattern Underneath All of Them

Every category above shares the same shape: the agent does the assembly, and a human owns the decision.

That isn’t a temporary limitation to be engineered away in the next release. For anything touching clinical judgment, it’s the design target. The regulatory direction across agencies has been consistent, AI should augment clinical judgment rather than replace it.

Teams that treat human-in-the-loop as a compliance checkbox to be minimized tend to build systems that are harder to deploy and harder to defend. Teams that treat it as a core architectural constraint tend to ship. Our eBook on AI Agents & CDS Hooks goes deeper on wiring these decision points into live clinical workflows.

Agent use cases by autonomy level and governance burden (illustrative)
Figure 4 Agent use cases by autonomy level and governance burden (illustrative)

V. A Practical Sequence for Teams Building Now

If you’re a product or engineering team looking at the gap and trying to figure out where to start, a rough order of operations:

1. Pick a Workflow Where Writes Are Reversible

Not “low risk” in the abstract, specifically reversible. If the worst outcome is a resubmitted form rather than an incorrect medication change, you have room to learn in production.

2. Build the Permission and Audit Layer Before the Agent

This inverts how most teams sequence the work, and it’s the single highest-leverage change. The agent is the easy part now. The delegation model, the scoped grants, and the decision-level audit trail are what determine whether it can ever leave the pilot.

3. Define the Escalation Path Before the Happy Path

For every agent, write down: what conditions trigger human escalation, who receives it, what the response SLA is, and what the agent does while waiting. If you can’t answer these, the agent isn’t ready regardless of how well it performs.

4. Build an Omission Test Set

Alongside your accuracy evaluation, build a set of cases where the correct action is to stop, flag, or escalate. Measure the miss rate. Treat regressions here as blocking.

5. Stage the Rollout With Hard Caps

One site, then one region. Daily action limits. A kill switch that a non-engineer can operate. Remove the constraints as evidence accumulates, not as confidence does.

How Mindbowser Can Help

Most of the work in closing the 69/22 gap happens below the model layer, in identity, integration, permissions, and audit. That’s the layer we’ve been building in healthcare for over a decade.

1. FHIR-Native Data and Access Architecture

Agents are only as good as the data layer they sit on. Our ConnectHealth platform provides pre-built HL7 and FHIR integration with Epic, Cerner, Athenahealth, and NextGen, with normalized patient data and infrastructure designed for controlled PHI access, which is the foundation any agent deployment needs before it needs a model.

2. Agent Design With Governance Built In

Our AI agent accelerators, including InsureVerify AI for eligibility verification, MedAdhere AI for adherence monitoring, RPMCheck AI for remote monitoring check-ins, and DischargeFollow AI for post-discharge follow-up, were built around the constraint set described in this article: scoped access, bounded autonomy, and explicit escalation paths to clinical teams.

3. Compliance-First Engineering

We build to HIPAA, SOC 2 Type II, GDPR, and FDA SaMD guidelines, with security integrated into the SDLC rather than reviewed at the end. For agent projects specifically, that means audit design and permission modeling happen during architecture, not during remediation.

4. Clinical Validation and Workflow Design

Our teams pair engineering with clinical subject matter expertise, which matters most for the questions that don’t have technical answers: where the human gate belongs, what should trigger escalation, and how to design a review step that clinicians actually perform rather than click past.

Planning an agent deployment and want a second opinion on the architecture?

Final Thoughts

The 69/22 gap is often framed as a story about caution, healthcare being slow, conservative, resistant to change. That framing is flattering to vendors and unhelpful to builders.

The more accurate story is that healthcare organizations adopted the AI that fit their existing infrastructure, and haven’t yet adopted the AI that requires new infrastructure. Generative AI slotted into a read-only, human-decides model that hospitals already knew how to govern. Agents don’t slot in anywhere. They require an access model that most healthcare systems have never had to build, because until now nothing but a person ever wrote to the chart.

That work is unglamorous and it’s most of the job. But it’s also the durable part. Models will keep changing. The identity, permission, and audit layer you build underneath them is what makes the next model deployable in weeks instead of quarters.

The organizations that close the gap first won’t be the ones with the best model access. They’ll be the ones that built the substrate.

Amreetanshu Kumar Sinha

Amreetanshu Kumar Sinha

Trainee Software Engineer

Connect Now

Amreetanshu is a full stack developer at the start of his software engineering journey, working across the MERN stack and Java. He has contributed to multiple healthcare projects, where he has learned to build applications that balance clean functionality with the demands of real-world clinical workflows. Alongside development, he brings an eye for design and product thinking, enjoying the process of shaping features from initial idea to working solution. Curious and driven, he is committed to deepening his expertise and growing into a well-rounded engineer who builds technology that makes a meaningful difference.

Share This Blog

Read More Similar Blogs

Let’s #Transform Healthcare,# Together.

Partner with us to design, build, and scale digital solutions that drive better outcomes.

Location

Global Tech Teams LLC, 525 Washington Blvd, Industrious at Newport Tower, Jersey City, NJ 07310, United States.

Contact

+1 408 786 5974
contact@mindbowser.com
BOOK A QUICK CONSULTATION

Have a Healthcare Project in Mind?

Let’s discuss your goals, workflows, and next steps in a focused consultation call.

Calendar icon Schedule a Call

Contact form