Has AI really surpassed physicians at cognitive medical tasks?

Share
Has AI really surpassed physicians at cognitive medical tasks?

Last week my feed was full of discussion about this JAMA perspective piece, which argues that AI won’t just assist physicians, it may soon replace them.

The authors argue that for five core cognitive medical functions, the literature shows that AI is already as good or superior to physicians, and that AI alone will soon be the highest standard of care for these tasks. Those core functions are:

  • eliciting medically relevant information (taking a patient history)
  • making a differential diagnosis
  • selecting testing
  • prescribing treatment
  • managing chronic disease

The perspective claims to have reviewed “all published articles on AI in medicine since January 1, 2024” and found that “already generative AI rivals or outperforms licensed physicians” at these tasks.

What struck me about all the conversation around this piece is that most people seemed to take this claim at face value — that the literature really does show this. Criticism centered around whether or not AI should replace physicians even if it can, whether the scenarios in the literature really reflect real-life patient care, and potential conflicts of interest by the authors (one owns a health company that uses AI, and another is his billionaire father).

But something else about the article (mainly, a total lack of methods) gave me pause, and once I took a quick look under the hood of the literature review, things started to unravel a bit. Let’s take a look…

No methods section

The first thing that caught my attention was a total absence of a methods section. This article was unusual in that it was a perspective piece, but it reported having done a literature review. Perspective pieces don’t require a methods section, but comprehensive literature reviews generally do — usually details about how the authors conducted the review are included to make sure it is 1. reproducible and 2. didn’t introduce a serious bias or miss a bunch of articles. Standard details normally included describe which databases were searched, which specific search terms were used, study inclusion and exclusion criteria, and more.

This perspective piece includes none of these details — only a single sentence is provided:

“Review of all published articles on AI in medicine since January 1, 2024, shows that medicine is rapidly approaching the transition point at which AI alone will exceed physicians and physician-AI hybrids in providing the best care at 5 fundamental cognitive medical tasks.”

Again, as a perspective piece, a methods section is not formally required. But it’s odd to not have any information — at all — about how the literature search was done.

Do the studies actually support the claims?

Next let’s look at the results of the literature review, which are summarized in a very long supplemental file, listing dozens of studies.

It’s a long list, so let’s focus on the first medical task that AI can allegedly do as well as physicians: taking a patient history. (The patient history is the conversation you have with your doctor where they ask you why you came in, what your symptoms are, review your medical history, etc.)

Three studies are listed in this section, which is labeled “AI alone exceeds physicians alone” at “eliciting patient past medical history and other information.” Based on these labels, I would assume these three studies:

  1. tested if AI could elicit a history from a patient (i.e. have the conversation),
  2. compared that to histories taken by physicians, and
  3. found the AI history was better.

Let’s see if that’s actually true.

Article #1: Privacy-preserving large language models for structured medical information retrieval

(Wiest, I et al., NPJ Digit Med 2024)

What this study did: This study used AI to analyze patient medical records. Specifically, it was fed free text documented by clinicians, and the AI model was asked to figure out when the clinician had documented abdominal pain, ascites, shortness of breath, confusion, or liver cirrhosis. AI did not interact with patients; it was fed information documented in the medical chart.

Did this study test if AI could elicit medically relevant information from a patient? Not at all.


Article #2: Adapted large language models can outperform medical experts in clinical text summarization

(Van Veen, D et al Nat Med. 2024)

What this study did: This study asked clinicians and AI to summarize medical records. Medical records can include hundreds of clinical notes, radiology reports, lab tests, and more, and summarizing these to get a high-level view of what’s going on with the patient can be a giant task. The study compared AI- and physician-generated summaries, and found that the AI summaries were preferred 36% of the time, physician summaries were preferred 19% of the time, and the rest were rated as no different. Again, AI did not interact with the patients in this study. It was fed the medical records.

Did this study test if AI could elicit medically relevant information from a patient? Not even a little bit.


So far we are 0 for 2! Neither of these studies showed anything close to what is claimed in the perspective piece — that AI is as good or better than doctors at taking a patient history. Neither study even attempted to have AI elicit information from the patient. It is weird (and methodologically inappropriate!) to have presented these studies as evidence that AI can now beat doctors at this task. That is not at all what they show.

Now, for the last study…

Article #3: Towards conversational diagnostic artificial intelligence

(Tu, T et al Nature. 2025)

What this study did: This is the only study that actually had AI take a history directly from a patient, but with some big caveats. It’s an interesting study: they created a simulated environment and randomized patient actors to either 1. interact with AI or 2. interact with a primary care doctor. Because it was randomized and blinded, all interactions were done via text, meaning the doctors could only type their questions for their patient (“what brought you in today?”), and the patient actors typed back their responses. The study found that AI had higher diagnostic accuracy and was rated as having higher conversational quality than the primary care physicians (possibly in part due to longer responses provided by AI in this time-constrained setting.)

This study is impressive, but does it show that AI is on it’s way to replacing physicians when it comes to taking a medical history?

Not really. There are like 97 different things we could point out about how this simulation doesn’t translate to actual patient care across medical disciplines. All interactions in the study were done over text, putting the doctors at a disadvantage (they had to pretend to be a chatbot, which is not how physicians take histories). It used standardized cases with patient actors, which doesn’t capture the messiness of real-world patient scenarios. None of the patients were drunk, or refused to answer a question, or repeatedly went off on tangents, or forgot what you were asking them. None of them were blind, or unable to hear, or unable to type, or forgot why they came to see you, or forgot where they were at all. None of them were delirious, or having a stroke, or were uncontrollably crying, or were so critically ill they could barely breathe. This is the reality of caring for patients in the real world: when people are sick, they can’t always neatly type their answers.


That was all the studies cited to back the claim that “AI gathers patient information as effectively as, if not more effectively than, physicians.” So what do we have? Two studies that don’t even test this question, and a third that is interesting, but very preliminary. None of these convincingly show that AI is so good at taking patient histories that it will soon replace physicians in the real world.

One more note on this literature review: it was also missing a highly relevant study for this section. I am again left wondering how this literature search was done, and what tools may have been used.

The thing AI will never be able to do

To end, I will join the chorus of criticism reminding people that these cognitive medical tasks are not the entirety of practicing medicine. We could talk about the physical exam (so analog!), procedures, surgery, or what it would be like to have a chatbot break the news that a love one has died.

But suppose AI gets so good it can do all that too. There is one thing it will never be able to replicate: a human conscience that has a duty to take care of you. This may sound trite, but pause and think about a world where there is no human barrier between you and the for-profit company providing your medical care. While physicians are far from perfect humans, the vast majority of them take their duty to their patients seriously, even when it means pushing back against the financial pressures within medical systems that are causing their patients harm.

Now imagine that your doctor is that for-profit system. An algorithm that can be slightly tweaked to avoid diagnosing an expensive condition, or “forget” to ask you a question that will lead to an expensive treatment. If AI actually replaced physicians, the potential for financial exploitation would be huge, all hidden within the black box of the LLM. Human physicians are imperfect, but they are also fighting back against a system that is so often failing (and bankrupting) patients. AI may get high marks on empathy, but it has no conscience that compels it to do right by you.


Kristen Panthagani, MD, PhD, is completing a combined emergency medicine residency and research fellowship focusing on health literacy and communication. In her free time, she is the creator of the newsletters You Can Know Things and The Public Health Roundup. You can also find her on Instagram, Threads, and LinkedIn. Views expressed belong to KP, not her employer.