General-Purpose LLMs Outperform Specialist Clinical AI Tools and Physicians in Benchmark Studies
Jul 31, 2026
Benchmark studies show general-purpose LLMs like GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 outperform FDA-cleared clinical AI tools and physicians on diagnostic reasoning tasks. Domain-adapted LLMs also beat larger generic models on clinical research tasks, while agentic systems show only modest gains. Regulatory clearance pathways have not yet addressed comparative performance.