Researchers Tested AI as a Real Healthcare Worker — and It Fell Short

News
Friday, 03 July 2026 at 21:17
Onderzoekers testen AI zoals een echte zorgmedewerker en de resultaten vallen tegen
A new scientific benchmark shows that even the most advanced AI agents still struggle with end-to-end care workflows. Even Codex GPT-5.5, the top performer in the test, successfully completes only about 42 percent of medical tasks. The researchers stress that while AI is improving, it remains far from fully autonomous use in healthcare.
The benchmark, HealthAgentBench, was published on June 30 on arXiv by researchers from Microsoft Research and several international institutions. Instead of quizzing AI with isolated medical questions or multiple-choice exams, the benchmark targets realistic clinical workflows where AI agents must independently execute multiple steps.
That makes the results far more relevant to real-world practice.

Why this benchmark stands out

Most AI tests measure how well a language model answers individual questions.
HealthAgentBench takes a completely different approach.
The benchmark comprises 54 realistic care tasks across seven categories. AI agents must operate within real healthcare environments, draw from diverse data sources, and handle complex software. They receive only limited instructions and must independently decide which steps are needed to complete a task successfully.
Among other things, the researchers test:
  • electronic health records (EHRs)
  • medical imaging
  • clinical scheduling
  • scientific research
  • data analysis
  • diagnostic workflows
  • complex multi-step decision-making
So this isn’t a chatbot answering a question—it’s an AI agent attempting to run a full medical workflow.

Even GPT-5.5 stalls at roughly 42 percent

And that’s exactly where current AI systems hit hard limits.
According to the researchers, Codex GPT-5.5 earns the highest overall score across all models tested, yet successfully completes only about 42 percent of full tasks.
Other frontier models perform even lower.
That doesn’t mean AI makes medical errors 58 percent of the time. The benchmark only marks a task as successful if the entire workflow is completed correctly. If an AI agent stalls at any point or skips a critical step, the attempt is counted as a failure.
That yields a more realistic picture of what AI can truly do on its own today.

Medical imaging remains a major roadblock

The researchers see clear differences across types of medical work.
AI agents do relatively well at automatically building research pipelines from EHR data.
It gets much harder when multiple skills are required at once.
Medical imaging, in particular, is a persistent challenge. AI must not only interpret images, but also synthesize information, plan follow-up actions, and make decisions within a complex clinical context.
Tasks that require searching large volumes of information or chaining many reasoning steps also remain difficult for all current models.

Why this matters for the AI industry

The timing is striking.
More tech companies are pitching AI agents as the next big leap in artificial intelligence. Instead of just generating text, these systems are supposed to run complete processes, operate software, and make decisions autonomously.
Healthcare is often cited as one of the most promising use cases.
But this benchmark shows the leap from a smart chatbot to a dependable digital coworker is far bigger than many suggest.
An AI agent needs not only medical knowledge, but also the ability to handle diverse data sources, software environments, edge cases, and complex decision-making. For now, that remains a far tougher challenge.

The warning extends beyond healthcare

While HealthAgentBench was designed for healthcare, its conclusions reach far beyond hospitals.
Nearly every sector is experimenting with AI agents to automate administrative processes, analyses, or business workflows.
The benchmark shows those long chains of actions are still a weak spot. A model may execute single instructions well but loses the thread when it has to make dozens of decisions in sequence.
That’s a crucial signal for companies expecting AI agents to take over entire business processes in the near term.

A more realistic yardstick for next-gen AI

The researchers present HealthAgentBench as a new standard for measuring progress in AI agents.
Where existing benchmarks are often saturated—modern models score nearly perfectly—this test exposes where the biggest challenges still lie.
That makes the benchmark useful for developers, healthcare providers, and policymakers assessing when AI is truly ready for large-scale deployment in critical environments.
The takeaway is clear: despite rapid progress, today’s best AI agents still lack the reliability needed to independently execute complex medical workflows. There’s significant room for improvement before AI can serve as a fully-fledged digital colleague in healthcare.
loading

Latest comments

    Loading