Every few weeks, another headline announces that artificial intelligence has matched a doctor, beaten a doctor, or is finally ready to become one. The latest wave arrived with familiar force: Several studies claim that AI outperforms physicians on clinical reasoning tasks.

The implication is hard to miss. Doctors are slow, expensive, and human. AI is faster, and potentially better.

Headlines about AI replacing doctors have already helped fuel a parallel, industry-driven medical system outside the walls of the hospital.

More than 40 million Americans ask ChatGPT a health question every day. Most of those conversations happen outside clinic hours, and most of them lead not to a doctor but back to the user — someone with no training in what to provide, how to prompt, or how to critically appraise what comes back. The models themselves are doing less and less to redirect them to appropriate care pathways. A study published this year found that medical disclaimers, once standard in chatbot answers to health questions, have largely disappeared; today’s leading models will not only respond to health questions but ask follow-ups and attempt a diagnosis.

Patients can act on the answers without ever seeing a doctor. Oura sells a 50-biomarker blood panel through Quest Diagnostics for $99. Function Health, valued at $2.5 billion in November, lets members order 160 lab tests a year, schedule a full-body MRI, and authorize ChatGPT to read the results. Ro and Hims will write prescriptions for weight loss or anxiety after an asynchronous intake. Doctronic, which calls itself “the world’s #1 AI doctor,” has run 24 million consultations and now writes AI-generated prescription refills in Utah. The work that used to happen inside a clinic is becoming a set of consumer products that patients can assemble on their own. The era of true patient empowerment of data, disease, and management in health care has never been closer.

Why is this transformation occurring at such speed and scale in medicine, rather than in law, finance, or software, where AI is also being used as an accelerant? Money. Health care represents close to one-fifth of the American economy. Capturing part of a doctor’s work, or persuading a hospital, an insurer, or a patient that they can save money without sacrificing quality of care, represents an enormous financial incentive. This has driven every major AI company into health care. Headlines implying that AI is ready to replace doctors only fuel the market.

Recent studies have shown that advanced AI systems can perform impressively on selected clinical tasks. At the same time, researchers have begun to ask whether autonomous clinical AI should be licensed and regulated more like clinicians than conventional medical devices. That debate is necessary. It also reveals the central risk of this moment: Evidence generated in bounded research settings is being pulled into a marketplace eager to claim that doctors are becoming optional. A model that performs well on isolated clinical problems may be useful, even transformative. But it does not follow that it should operate as a substitute for clinical judgment or responsibility.

Beyond the headlines lies a more basic confusion that even most researchers fail to grasp — clinical reasoning and model reasoning are not the same thing. The hidden clinical reasoning process in medicine is one that you will not find published online or in books, and certainly not in the corpus of data AI trains on. It includes mental musings while reading the initial triage note, abandoned hypotheses, and nuanced decision-making — all forks in the road and potential off-ramps that an observer never sees. Navigating these forks is a mental burden that separates experienced physicians from those just starting out. AI not only lacks this hidden framework and data repository for learning, but it arrives at conclusions using a fundamentally different method.

What the models are doing instead is genuinely unclear. Anthropic, OpenAI, and Google have all built research programs aimed at understanding how their own systems arrive at answers, and the field of mechanistic interpretability — understanding how neural networks work internally — exists because the engineers who built these models cannot fully explain them. Interpretability remains a nascent, research-oriented field, especially when it comes to clinical concepts and decisions. That should give pause to anyone building a medical product around these models.

This matters because the moment patients turn to one of these tools is, almost by definition, a moment of uncertainty and anxiety for them. To understand how these systems perform across the steps of a clinical encounter, we recently tested 21 frontier models in a study published this year in JAMA Network Open — the largest evaluation of these systems across the full arc of clinical reasoning to date. Given a complete case, they named the correct diagnosis more than 90% of the time. Given only what a clinician would gather at the start of a visit, they failed to produce a comprehensive differential more than 80% of the time. Medicine is high-stakes: fail to consider all possibilities, and you may never make the right diagnosis.

Our results almost certainly overstate how well these tools perform outside the clinic. At home, the model is working from a different kind of input: scattered symptoms as they are experienced, no physical exam, and no one, let alone an experienced health care professional, separating signal from noise. The model is being asked to do precisely the part of medicine it does worst. The gap from initial triage to final diagnosis is the same one prone to severe, preventable errors. Go down the wrong path, and you might not get to the right one in time.

These problems converge on a question of responsibility. Many studies and industry claims evaluate AI by asking whether it is “non-inferior to physicians.” This sounds like a careful threshold, but it is the wrong one for deciding whether a system should assume a clinical role. A physician’s job does not end with the diagnosis. It includes an obligation to remain accountable when a wrong decision harms a patient. Current AI systems assume no comparable obligation, and imposing one would create substantial financial and legal liabilities that the companies building them have little incentive to accept. When a physician is wrong, accountability is imperfect but identifiable. When an AI system is wrong, responsibility disperses and, in practice, will often still fall on the clinician.

We study and build upon these tools because we believe they hold enormous potential in medicine — our laboratory published some of the earliest research on large language models (LLMs) for clinical decision support, and continues to lead investigations in this space. Used well, AI can deepen and extend the doctor-patient relationship by helping patients understand their conditions, stay connected between visits, and integrate the longitudinal data that a clinician alone cannot reasonably synthesize in a 15-minute appointment. It can take on much of the administrative work that has steadily pulled doctors away from their patients. In this way, AI can serve not as a substitute for medicine, but as a tool that can help medicine become its best version — more attentive, more human.

That is not the system being built by most of the companies riding the AI wave toward their next funding round, many of them little more than a polished user interface wrapped around someone else’s model. Patients are not wrong to look for alternatives in an oversubscribed, underfunded health care system. But they should not be guided into a shadow system that borrows the authority of medicine while avoiding its responsibilities.

IBM put it best in a 1979 internal training manual: “A computer can never be held accountable, therefore a computer must never make a management decision.”

Arya Rao is an M.D.-Ph.D. candidate at Harvard Medical School and MIT, and a Paul and Daisy Soros Fellow. Marc Succi, M.D., is an associate professor of radiology at Harvard Medical School, inventor, and radiologist at Mass General Brigham.