General-Purpose LLMs Outperform Specialist Clinical AI Tools and Physicians in Benchmark Studies

Benchmark studies show general-purpose LLMs like GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperform FDA-cleared clinical AI tools and physicians on diagnostic reasoning tasks. Domain-adapted LLMs also beat larger generic models on clinical research tasks, while agentic systems show only modest gains. Regulatory clearance pathways have not yet addressed comparative performance.

Head-to-head studies published in Nature Medicine and Science show that general-purpose large language models (LLMs) outperform both FDA-cleared clinical AI tools and physicians on case-based diagnostic and clinical reasoning tasks. In a Nature Medicine benchmark published June 23, 2026, OpenAI's GPT-5.2, Google's Gemini 3.1 Pro, and Anthropic's Claude Opus 4.6 outperformed two specialist clinical AI tools, OpenEvidence and Wolters Kluwer's UpToDate Expert AI, on questions submitted by practicing physicians. The general-purpose models won across every medical benchmark tested, and neither cleared tool held a performance advantage on real-world, unstructured point-of-care queries.

A separate study published in Science found that a large language model from OpenAI can outperform physicians in case-based diagnostic and clinical reasoning evaluations, including one experiment using real-world data from a Boston emergency department. The paper's co-senior author expressed worry that the experiments, all based on simulated and historical cases, could be misconstrued as proof of AI's safety and efficacy when used to treat real patients.

The Nature Medicine results align with an earlier July 2024 study in the Journal of Medical Internet Research, in which ChatGPT with GPT-4 outperformed emergency department resident physicians on diagnostic accuracy across 100 internal medicine emergency cases, achieving superior accuracy compared with both GPT-3.5 and the human physicians. A resource called TrialPanorama, described in npj Digital Medicine, aggregated 1.6 million clinical trial records from fifteen global registries and constructed 152K training and testing samples spanning eight clinical research tasks. Benchmarking cutting-edge LLMs on these tasks revealed limited clinical reasoning capability in generic LLMs, but an 8B LLM developed on TrialPanorama using supervised fine-tuning and reinforcement learning outperformed 70B generic counterparts across all eight tasks, with relative improvements of 73.7%, 67.6%, 38.4%, 37.8%, 26.5%, 20.7%, 20.0%, 18.1%, and 5.2%.

Benchmarking of agentic AI systems for clinical decision tasks found more modest gains. A study evaluated the open-source OpenManus, built on Meta's Llama-4 and extended with medically customized agents, and Manus, a proprietary agent system employing a multistep planner-executor-verifier architecture, across AgentClinic, MedAgentsBench, and Humanity's Last Exam (HLE). Despite access to tools such as web browsing and code execution, the systems reached 60.3% and 28.0% in AgentClinic MedQA and MIMIC, 30.3% on MedAgentsBench, and 8.6% on HLE text, with multimodal accuracy as low as 15.5% on multimodal HLE and 29.2% on AgentClinic NEJM. Resource demands increased substantially, with more than 10× token usage and more than 2× latency, and although 89.9% of hallucinations were filtered by in-agent safeguards, hallucinations remained prevalent.

The performance gap between cleared and uncleared tools raises regulatory questions. A cross-sectional study analyzed in JAMA Health Forum examined 691 FDA-cleared AI and machine learning devices cleared between September 1995 and July 2023; only 1.6% of those devices reported data from randomized controlled trials in their benefit-risk documentation prior to clearance, and fewer than 30% shared key safety and adverse event information before approval. The FDA issued revised Clinical Decision Support software guidance on January 29, 2026, replacing the prior version from September 28, 2022, and its December 4, 2024 final guidance on Predetermined Change Control Plans (PCCPs) for AI-enabled device software functions requires sponsors to pre-specify types of changes and testing protocols. A February 10, 2025 warning letter to Exer Labs for its Exer Scan AI diagnostic tool illustrates enforcement focused on tools marketed without clearance.

Adoption of AI in clinical practice is already widespread: the AMA's 2024 physician AI survey found 66% of physicians currently use AI in practice, up from 38% in 2023, and 68% reported perceiving definite advantage from AI tools. The EMA's 2024 reflection paper on AI/ML in clinical development adopted a risk-tiered approach, with high-risk tools that could influence trial primary endpoints or safety reporting facing heightened scrutiny, including prospective performance monitoring requirements. The Tufts Center for the Study of Drug Development's 2024 global AI adoption assessment, conducted with the Drug Information Association between May and August 2024, found that adoption of AI-enabled solutions in clinical development is accelerating rapidly despite the absence of standardized performance benchmarks across vendors.

Related Entities

Related Articles

References

  1. Benchmarking and developing large language models using one million clinical trials · nature.com
  2. Nature Medicine's June 2026 Benchmark Study Reveals General-Purpose LLMs ... · clinicaltrialvanguard.com
  3. The Specialized Clinical AI Trap: General-Purpose Chatbots Are Winning the Benchmark War · clinicaltrialvanguard.com
  4. Clinical large language model centered on electronic medical records | npj Digital Medicine · nature.com
  5. Large Language Model–Generated Patient Instructions for Prescriptions in Primary Health Care · jmir.org
  6. As artificial intelligence show off diagnostic chops, scientists reckon with the way forward · statnews.com
  7. Benchmarking large language model-based agent systems for clinical decision tasks · nature.com