As AI Enters Clinical Practice, Researchers and Medical Groups Urge Boundaries
New research and guidelines highlight the promise and perils of AI in medicine, from a Dartmouth study showing AI-drafted patient messages can introduce errors, to the Israel Medical Association's new position paper. Regulatory responses vary, with FDA guidance and state laws still evolving.
Artificial intelligence is moving into emergency rooms, community clinics, and patient portals, but a growing body of research and professional guidance warns that safeguards are needed. A Dartmouth study presented at the 2026 Annual Meeting of the Association for Computational Linguistics found that AI-drafted responses to patients can introduce errors and extraneous details, while the Israel Medical Association has published a position paper defining boundaries for the technology's use in medicine.
The Dartmouth researchers analyzed 146,000 conversations between 10,105 patients and their primary care physicians at a large rural health system, in the first large-scale study of an online patient portal that uses AI to draft responses. They found that AI-generated answers frequently misalign with what clinicians would actually write — responses that are too long, don't ask follow-up questions, and use irrelevant or inaccurate medical details. As a result, physicians may spend more time editing responses than it would have taken to write them, the researchers report. In one example, the portal's AI suggested telling a 32-year-old woman taking an acid reflux drug and concerned about constant nausea that the medication might require a diet adjustment; a physician instead asked if there was any chance she was pregnant.
The researchers evaluated responses drafted by Claude, Gemini, and ChatGPT, as well as three smaller commercial platforms, Llama, Aloe, and Qwen. They showed that adapting AI to how individual physicians communicate can improve accuracy by 33% and reduce editing by 26%. The team also developed a technique called TADPOLE — Thematic Agentic Direct Preference Optimization for Learning Enhancement — that trains AI platforms to better match physicians' standards for precision and information quality. Plugging TADPOLE into the six commercial LLMs produced drafted responses that better matched those standards, which one co-author said could save a busy clinician an hour or two of work a day.
The Israel Medical Association, through the Institute for Quality in Medicine and the Israeli Society for Risk Management and Patient Safety in Medicine, published a position paper this month that attempts to define the rules of the game: to use the technology but not become enslaved by it; to be assisted by it to improve diagnosis and treatment, but not turn it into a hidden doctor making decisions. The paper notes that as of May 2024, the US Food and Drug Administration had approved 882 medical devices that use artificial intelligence, with radiology accounting for 671 devices, or 76%. It also describes systems such as Rounds, which records a medical visit, transcribes the conversation, and generates a full visit document, and Clalit Health Services' AI-PRO, which scans medical records nightly and surfaces recommendations and at-risk patients to family doctors.
The position paper emphasizes that the promise is great but so are the risks. An AI system's recommendations can be biased if the training data is partial, biased, or unrepresentative of the patient population. Another risk is a false sense of security: when a computerized system gives a fast, well-phrased, and self-confident answer, it is easy to believe it.
The regulatory environment is fragmented. In January, the FDA updated its software guidance to allow AI tools to operate with less oversight when assisting doctors, placing software that enables a physician to independently review the basis for an AI recommendation outside the agency's regulation of medical devices. However, the carve-out only covers AI with a doctor in the loop; there is no comparable exemption for AI that talks directly to patients, making the most autonomous clinical AI the least regulated. States have moved in different directions: Utah, Arizona, and Texas are building frameworks to accelerate deployment, while New York and California are moving to curtail AI in medicine. New York lawmakers introduced legislation that would render it potentially illegal for AI to provide even basic medical guidance, and California enacted a law mandating disclosure to patients when AI is involved.
Meanwhile, data shows that one in three Americans are now turning to AI chatbots to diagnose symptoms and direct care, a figure that doubled in a single year. Some physicians see potential for AI to expand access. A pediatrician has argued that a well-designed AI could simulate 'serve and return' interactions for children, asking questions and adjusting based on responses, unlike passive screen content. The federal government is also soliciting proposals to develop AI that will independently manage heart failure events, a disease for which only 1% of patients receive the recommended medication regimen and five-year mortality rates exceed 50%.