Kristen, this is an excellent and much needed breakdown. I would not have looked at the paper this closely without your post.
The study is interesting and clearly shows strong performance of LLMs in generating differential diagnoses from curated clinical data. But it is important to be precise about what was actually tested. The ER component evaluates a second opinion diagnostic task using physician generated information, not real time triage.
In practice, triage is centered on prioritization and identification of dangerous conditions rather than diagnostic completeness. It is also a systems process, often initiated by nursing assessment and protocol driven decisions, not a standalone diagnostic exercise. That layer of decision making is not directly modeled or evaluated here, which limits how well these results translate to actual emergency care. It also helps explain why headlines suggesting that AI has “outperformed ER doctors at triage” are misleading.
Thank you so much for this! It's exactly the kind of critical lens this space needs.
I'm coming at this as a layman and as the President of a company that builds AI for emergency department triage, Mednition (KATE AI), so I'm neither a clinician nor a disinterested observer. But your breakdown of what this study actually tested and what the headlines got wrong resonates with what we try to communicate to health systems every day. It's a major challenge!
The point you make about ER triage being fundamentally about not missing the emergency rather than generating a comprehensive final diagnosis is one we've built our entire clinical philosophy around. General-purpose LLMs fed a completed medical record is a very different thing from an AI system designed to support a triage nurse's real-time acuity decision on a patient walking through the door. The study conflates the two, and the media ran even further from there.
Your concern about public overconfidence in AI diagnostics is one our team shares, frankly, from the supply side. Headlines like these make it harder, not easier, to have honest conversations with clinical leaders and executive about what AI tools can and cannot responsibly do in the ED.
The hype creates a credibility gap that affects everyone in this space, including those of us trying to do it carefully.
We've now triaged over 5M encounters and growing. But it amazes me how many conversations we've had in recent months about LLM doing what our company researched for 5 years, used thousand of real patient records, and thousands of gold cases, peer reviewed and validated by a large team consisting of experienced ED physicians and nurses, review cases and our models on a weekly basis, give/receive feedback in real-time daily from practicing clinicians, leverage the ESI standard, have an exclusive partnership with the Emergency Nurses Association, etc. but have our credibility challenged while LLMs just skate by.
Really appreciate the clarity you bring here. Love to connect with you live some time. Subscribing.
I think there's also a naive fear that AI will replace doctors - or worse but better substantiated one that we'll see NP's plus AI viewed as a cost-effective replacement for doctors by investment firms.
As someone who has a weird pedigree (I was a software engineer for a lot of years before I went to med school, and I've lately been using AI as sort of a faster google to help me generate the odd code example as I'm developing a network-based system for veterinary blood banks), a lot of it has to do with what the model is fed as well. From what I've seen in the software world, ChatGPT and others are mostly scraping places like Reddit, which are littered with noob questions and bad advice, and then they collage these things uncritically into answers that are sometimes helpful, sometimes not, but always need to be double-checked. As an experienced clinician I can do the job far better and far faster without AI being "helpful." I'm neither a technophobe nor a tyro, but my experience as a clinician and as an engineer has led me to the conclusion that while AI is promising, much of the hype is extremely overblown.
That said, it's not just AI. I've seen people uncritically worship UpToDate, only to show them why UpToDate can't always be trusted. They used to be very uncritical about what they'd reference, and while it's gotten better, it too needs to be read with a critical eye.
Where AI would be most useful to emergency physicians is not in generating a differential, but rather in taking over some of the more mundane tasks. Charting is an excellent example. There is lots of work put into the non-doctoring aspect of charting to be sure you meet regulatory hoops, hit all the compliance things, etc. It's tedium that AI could help do really well, and the assist would be welcome. I'd still have to review and sign off on the chart, but having something assemble the pieces might be helpful.
Larry, really appreciate this perspective, especially coming from someone with your background on both sides of the stack.
Your point about documentation is actually where a lot of the clinical AI investment is happening right now with ambient documentation tools like Aridge, Nabla, and DAX Copilot which have lots of momentum. They listen to the patient encounter and draft the note in real time, so the physician reviews and signs rather than types from scratch.
Chart summarization tools are tackling the other end, synthesizing prior records so a clinician walks into the room already oriented rather than hunting through years of notes.
None of that replaces clinical judgment. It just gives it back more time. Which is, as you said, where the assist is actually welcome. Now what gets done with that time is going to be the battle - health systems want the physicians to use the saved time to see more patients and the physicians want to spend more time with patients and improving cafre and outcomes versus seeing more.
At the end of the day, human-in-the-loop AI is critical for healthcare even moreso than other industries and use cases.
If those articles omitted most of the information you provided, they did the public a disservice. There is no substitute for a physician trained in emergency medicine.
A few of the articles mentioned only two physicians were included, but I don’t believe any acknowledged they weren’t emergency medicine trained or highlighted the other issues.
As an EM/CCM physician nearly 20 years out from residency, and a former academic, I first wanted to say thank you for such a nice write-up.
But the other thing that this study says - given that the data input was obtained from EM docs doing their jobs (they produced the original charts) - is that even an AI model of EM docs does better than IM docs at doing what it is that IM docs think we EM docs do. :)
NPR updated their headline and issued a correction! Not perfect, but progress. https://www.npr.org/2026/04/30/nx-s1-5804474/ai-doctors-openai-patient-care-diagnosis
TechCrunch also updated their headline and included commentary from this post! Glad to see news outlets hearing the feedback and correcting course. https://techcrunch.com/2026/05/03/in-harvard-study-ai-offered-more-accurate-diagnoses-than-emergency-room-doctors/
Kristen, this is an excellent and much needed breakdown. I would not have looked at the paper this closely without your post.
The study is interesting and clearly shows strong performance of LLMs in generating differential diagnoses from curated clinical data. But it is important to be precise about what was actually tested. The ER component evaluates a second opinion diagnostic task using physician generated information, not real time triage.
In practice, triage is centered on prioritization and identification of dangerous conditions rather than diagnostic completeness. It is also a systems process, often initiated by nursing assessment and protocol driven decisions, not a standalone diagnostic exercise. That layer of decision making is not directly modeled or evaluated here, which limits how well these results translate to actual emergency care. It also helps explain why headlines suggesting that AI has “outperformed ER doctors at triage” are misleading.
1000% agree!
Thank you so much for this! It's exactly the kind of critical lens this space needs.
I'm coming at this as a layman and as the President of a company that builds AI for emergency department triage, Mednition (KATE AI), so I'm neither a clinician nor a disinterested observer. But your breakdown of what this study actually tested and what the headlines got wrong resonates with what we try to communicate to health systems every day. It's a major challenge!
The point you make about ER triage being fundamentally about not missing the emergency rather than generating a comprehensive final diagnosis is one we've built our entire clinical philosophy around. General-purpose LLMs fed a completed medical record is a very different thing from an AI system designed to support a triage nurse's real-time acuity decision on a patient walking through the door. The study conflates the two, and the media ran even further from there.
Your concern about public overconfidence in AI diagnostics is one our team shares, frankly, from the supply side. Headlines like these make it harder, not easier, to have honest conversations with clinical leaders and executive about what AI tools can and cannot responsibly do in the ED.
The hype creates a credibility gap that affects everyone in this space, including those of us trying to do it carefully.
We've now triaged over 5M encounters and growing. But it amazes me how many conversations we've had in recent months about LLM doing what our company researched for 5 years, used thousand of real patient records, and thousands of gold cases, peer reviewed and validated by a large team consisting of experienced ED physicians and nurses, review cases and our models on a weekly basis, give/receive feedback in real-time daily from practicing clinicians, leverage the ESI standard, have an exclusive partnership with the Emergency Nurses Association, etc. but have our credibility challenged while LLMs just skate by.
Really appreciate the clarity you bring here. Love to connect with you live some time. Subscribing.
I think there's also a naive fear that AI will replace doctors - or worse but better substantiated one that we'll see NP's plus AI viewed as a cost-effective replacement for doctors by investment firms.
As someone who has a weird pedigree (I was a software engineer for a lot of years before I went to med school, and I've lately been using AI as sort of a faster google to help me generate the odd code example as I'm developing a network-based system for veterinary blood banks), a lot of it has to do with what the model is fed as well. From what I've seen in the software world, ChatGPT and others are mostly scraping places like Reddit, which are littered with noob questions and bad advice, and then they collage these things uncritically into answers that are sometimes helpful, sometimes not, but always need to be double-checked. As an experienced clinician I can do the job far better and far faster without AI being "helpful." I'm neither a technophobe nor a tyro, but my experience as a clinician and as an engineer has led me to the conclusion that while AI is promising, much of the hype is extremely overblown.
That said, it's not just AI. I've seen people uncritically worship UpToDate, only to show them why UpToDate can't always be trusted. They used to be very uncritical about what they'd reference, and while it's gotten better, it too needs to be read with a critical eye.
Where AI would be most useful to emergency physicians is not in generating a differential, but rather in taking over some of the more mundane tasks. Charting is an excellent example. There is lots of work put into the non-doctoring aspect of charting to be sure you meet regulatory hoops, hit all the compliance things, etc. It's tedium that AI could help do really well, and the assist would be welcome. I'd still have to review and sign off on the chart, but having something assemble the pieces might be helpful.
Larry, really appreciate this perspective, especially coming from someone with your background on both sides of the stack.
Your point about documentation is actually where a lot of the clinical AI investment is happening right now with ambient documentation tools like Aridge, Nabla, and DAX Copilot which have lots of momentum. They listen to the patient encounter and draft the note in real time, so the physician reviews and signs rather than types from scratch.
Chart summarization tools are tackling the other end, synthesizing prior records so a clinician walks into the room already oriented rather than hunting through years of notes.
None of that replaces clinical judgment. It just gives it back more time. Which is, as you said, where the assist is actually welcome. Now what gets done with that time is going to be the battle - health systems want the physicians to use the saved time to see more patients and the physicians want to spend more time with patients and improving cafre and outcomes versus seeing more.
At the end of the day, human-in-the-loop AI is critical for healthcare even moreso than other industries and use cases.
If those articles omitted most of the information you provided, they did the public a disservice. There is no substitute for a physician trained in emergency medicine.
A few of the articles mentioned only two physicians were included, but I don’t believe any acknowledged they weren’t emergency medicine trained or highlighted the other issues.
As an EM/CCM physician nearly 20 years out from residency, and a former academic, I first wanted to say thank you for such a nice write-up.
But the other thing that this study says - given that the data input was obtained from EM docs doing their jobs (they produced the original charts) - is that even an AI model of EM docs does better than IM docs at doing what it is that IM docs think we EM docs do. :)