AI Beat the Doctor. Now What?

By The Functional Medicine Report™
Originally published August 21, 2026 in Issue #006 of The Functional Medicine Report™
AI is outperforming physicians on some controlled clinical reasoning tasks. But new research reveals something more complicated: giving a doctor access to powerful AI does not automatically make the doctor better. The next clinical advantage may depend on how human and machine learn to reason together.
Key Takeaways
- Access to a strong model did not make physicians better. In a randomized trial of 50 physicians, median diagnostic-reasoning scores were 76% with GPT-4 and 74% with conventional resources — not a statistically significant difference. The standalone model scored 16 percentage points higher than the physicians using conventional resources.
- Changing the workflow did. Physicians who completed a 20-hour AI-literacy curriculum and then used an LLM scored 71.4%, against 42.6% for the conventional-resources group. A separate trial using a structured AI-first workflow scored 85%, against 75% for conventional resources.
- Influence runs both ways. When the AI saw the clinician’s assessment first, complete overlap between their diagnoses rose from 3% of cases to 48% — which undermines the idea of an independent AI second opinion. Systematically biased AI predictions reduced clinician accuracy by 11.3 percentage points, and image-based explanations did not significantly repair the damage.
- Benchmark performance is not a patient outcome. In a 2026 trial across 16 Kenyan primary-care facilities covering 103 clinical officers and 9,691 patients, an LLM embedded in the electronic record improved documentation quality. Treatment failure within 14 days did not differ significantly between groups.
- No trial has validated an AI-supported reasoning system specifically for functional medicine or complex chronic illness. The potential is substantial; the outcome evidence does not yet exist.
The Model Was Better. The Team Wasn’t.
In a randomized trial, 50 physicians were given up to six difficult diagnostic vignettes. Half could use GPT-4 alongside the conventional resources available to the other half.
The median diagnostic-reasoning score was 76% with GPT-4 and 74% with conventional resources. The two-point difference was not statistically significant.
Then the researchers tested the model by itself. In that exploratory comparison, the standalone LLM scored 16 percentage points higher than the physicians using conventional resources. (JAMA Network Open)
These were controlled diagnostic vignettes—not live patient encounters or evidence that AI provides better patient care. But the finding was still striking.
The model performed remarkably well. Giving doctors access to it did not make their average performance meaningfully better.
How is that possible?
Because capability does not simply pass from a tool to its user.
A model may generate a broad differential, identify relevant evidence and suggest the next diagnostic step. The clinician still has to decide when to ask, how to frame the problem, which information to include, how much influence to give the answer and what deserves further investigation.
Those choices shape the result.
A practitioner who forms an assessment before consulting AI will interact with the model differently from one who begins by asking it for an answer. A clinician who requests evidence that challenges the leading theory may receive more value than one who asks the system to justify it. A polished explanation can be inspected carefully—or mistaken for proof.
The same problem becomes more consequential in the complex cases familiar to functional medicine practitioners. The record may contain years of symptoms, several diagnoses, conventional and specialty laboratory findings, medications, supplements, environmental exposures, previous interventions and conflicting responses over time.
AI can generate connections across all of it. That does not tell the practitioner which connections matter.
The first trial exposed the gap at the center of this story: a powerful model can possess clinical reasoning capability that the clinician–AI team fails to use well.
Opening another tab is not a clinical workflow.
Something Changed When the Workflow Changed
Two trials published in 2026 produced much larger improvements. Both changed more than the clinician’s access to a model.
In one trial, 60 physicians in Pakistan completed a 20-hour AI-literacy curriculum covering the capabilities, limitations and appropriate use of large language models. After the training, they were randomized to use either conventional resources alone or conventional resources plus GPT-4o.
Among the 58 physicians who completed the trial, the group with LLM access scored 71.4%, compared with 42.6% for the conventional-resources group. (Nature Health)
The result is important, and so is its limitation.
Every participant received the same training, so the trial cannot tell us how much the curriculum itself contributed. It showed a large advantage for LLM access within a trained group. It did not prove the ideal way to train every practitioner.
A separate U.S. trial tested a more deliberately structured form of collaboration. Seventy physicians worked with conventional resources or a custom system that gathered clinician and AI assessments, compared them, identified agreement and disagreement and generated a critique combining both perspectives.
Physicians using the AI-first workflow scored 85%. Those using the AI-second workflow scored 82%. Physicians using conventional resources scored 75%. The difference between the two AI workflows was not statistically significant. (npj Digital Medicine)
Again, the caveat changes the meaning.
This was not a test of an ordinary commercial chatbot. It was a test of a purpose-built reasoning process, and the study did not include a standard-chatbot comparison arm.
It would be tempting to connect these studies into a neat progression: access failed, training worked and workflow solved the problem. The research does not let us do that.
The studies involved different countries, clinicians, model versions, case sets, interfaces and controls. They were not successive stages of one experiment.
They do point toward a more useful question:
How clinicians use AI may matter nearly as much as what the AI itself can do.
For functional medicine, that changes the adoption decision. A practice cannot assume that buying a powerful system creates a clinical advantage if every practitioner must invent an individual method for using it.
The more complex the case, the more important it becomes to define when independent reasoning should be preserved, when AI enters, how competing explanations are compared and who decides what happens next.
The advantage may not come from having the best model. It may come from building the best way to use it.
AI Can Make You Better. It Can Also Make You Worse.
Collaboration introduces another complication: the clinician and the AI can influence each other.
In the custom-workflow trial, researchers conducted a post hoc analysis of 58 matched cases. When the AI produced its assessment before seeing the clinician’s conclusions, all three of the clinician’s initial diagnoses appeared in the AI output in only 3% of cases.
When the AI received the clinician’s assessment first, complete overlap rose to 48%. (npj Digital Medicine)
The analysis was small, so it cannot tell us whether this happens consistently. But it exposes an important vulnerability in the idea of an “independent AI second opinion.”
A second opinion loses much of its value when one side already knows what the other thinks.
The influence runs in the opposite direction too.
In another randomized vignette study, clinicians reviewed cases involving acute respiratory failure. Standard AI predictions modestly improved their diagnostic accuracy. Systematically biased AI reduced it by 11.3 percentage points.
Adding image-based explanations did not significantly remove the damage caused by the biased predictions. (JAMA)
A wrong answer does not become safer because it arrives with a persuasive rationale.
Even the collaborative-workflow trial, which improved average performance, contained cases in which AI involvement made the score worse. In its AI-second arm, clinically actionable scores increased in 52 cases, remained unchanged in 83 and fell in 12. (npj Digital Medicine)
These were simulated cases rather than harmed patients. They still expose the danger of judging every clinical interaction by the average improvement.
This matters acutely in functional medicine because a working theory can shape how dozens of findings are interpreted. Once a practitioner begins viewing a case through a favored explanation—an exposure, an infection, a pathway or a laboratory pattern—the prompt itself may encourage AI to build a more sophisticated argument around that frame.
The output can sound expansive while remaining anchored.
A stronger workflow protects disagreement. It gives the clinician and the AI enough separation to catch something the other missed. It makes supporting and opposing evidence visible. It identifies missing information and preserves uncertainty instead of compressing the case into one smooth, confident conclusion.
Disagreement can be useful clinical information. A system that immediately blends two competing assessments may remove the friction that should trigger a closer look.
Human oversight remains essential. Its value depends on the independence and quality of the thinking being applied.
A human can be present and still be misled.
Winning the Test Isn’t the Same as Helping the Patient
The most dramatic AI headlines usually come from controlled tests.
One preprint evaluated MAI-DxO, a system designed to work through difficult diagnostic cases sequentially. On a benchmark constructed from New England Journal of Medicine clinicopathological conference cases, it reported 80% diagnostic accuracy, compared with a 20% average among participating generalist physicians.
That result demonstrates impressive performance under the benchmark’s rules. The physicians and AI worked under different resource conditions, and the encounters were simulated. As of August 14, 2026, the paper remained an arXiv preprint last revised on July 2, 2025. (arXiv)
It did not show that AI provides better bedside care or produces better outcomes for patients.
The distinction becomes clearer when the evidence is separated into three levels.

Benchmark
A benchmark asks whether a model can perform a defined task under constructed conditions.
It tells us something about technical capability. It cannot establish how clinicians will use the tool, how the tool will behave inside a practice or what ultimately happens to patients.
Controlled clinical vignette
A vignette trial asks how practitioners reason through simulated cases when assigned different resources or workflows.
It can reveal changes in diagnostic reasoning, susceptibility to bias and the influence of interface design. It cannot reproduce the full reality of incomplete records, physical examination, interruptions, patient preferences, uncertain follow-up and responsibility for the consequences.
Real patient care
A pragmatic trial introduces a system into ordinary clinical work and measures what happens in a defined patient population.
A 2026 trial in 16 Kenyan primary-care facilities moved the evidence closer to that standard. The study included 103 clinical officers and 9,691 patients. Clinicians using an LLM-based system embedded in the electronic record produced stronger documentation across several expert-rated measures, including the appropriateness of the recorded diagnosis, the completeness of the note and the treatment plan.
The primary outcome was treatment failure within 14 days. It did not differ significantly between the AI-assisted and control groups. (Nature Medicine)
Researchers identified no intervention-related serious adverse events or overall safety signal. The study did not establish safety equivalence or rule out uncommon harms. (Nature Medicine)
The Kenya study should not be reduced to an AI success or an AI failure. It shows why the outcome being measured matters.
A system can improve documentation without changing short-term treatment failure. It may improve a practitioner’s reasoning score without improving a patient’s health. It may also create benefits that a particular trial was not designed or powered to detect.
Better reasoning scores are not automatically better patient outcomes.
That distinction becomes even more important in chronic care, where meaningful results may unfold across months, several treatment decisions and changing patient behavior. A compelling explanation during one encounter cannot stand in for an improved clinical trajectory.
The broader evidence base remains young. A 2026 early-access systematic evidence map included 55 studies and registered trials and found heterogeneous results, inconsistent efficiency findings and important reporting gaps. The journal still identified the manuscript as an unedited early-access version as of August 14. (npj Digital Medicine)
AI has already shown that it can perform. Medicine still has to show when that performance improves care.
So What Does the New Clinical Expert Look Like?
The payoff is not a smaller role for the practitioner.
It is a more demanding one.
Clinical knowledge remains essential. But as these systems improve, retrieving and synthesizing information may become easier to reproduce. The harder skills are framing the case, judging evidence quality, ranking probabilities, recognizing missing information, incorporating physical findings and patient context, managing risk and deciding what deserves action.
The new clinical expert will know how to preserve independent thinking long enough for the human and the machine to actually challenge each other.
That practitioner will use AI to widen a differential without surrendering control of it. They will recognize when the system has caught a contradiction and when it has merely restated the original theory with more confidence. They will interrogate sources, challenge assumptions, reject weak recommendations and recognize when the evidence remains too uncertain to justify another test or intervention.
They will also understand their own effect on the system.
For functional medicine, this is an unusually demanding test. Practitioners may be reasoning across a long timeline, multiple body systems, conventional and specialty laboratories, medications, supplements, previous diagnoses, failed interventions, environmental factors and patient-reported responses over time.
AI may help hold that complexity in view. It may identify the result that does not fit, preserve a hypothesis that was prematurely discarded or surface a question nobody has asked.
It can also industrialize a familiar weakness: each abnormality gets a mechanism, each mechanism gets a product, and the plan grows faster than the clarity.
More data is not automatically better reasoning.
The value of AI in a functional medicine case should be judged by the quality of the questions it helps the practitioner answer:
What matters most? What does not fit? Which explanation has meaningful support? Where is the evidence weak? What information is missing? What should happen first?
The evidence reviewed for this story included no trial validating an AI-supported reasoning system specifically for functional medicine or complex chronic illness.
The potential is substantial. We do not yet have the clinical evidence to say that it improves outcomes in complex chronic care.
The field should demand evidence that AI improves prioritization rather than merely increasing the number of connections, tests and protocols a practitioner can generate.
Six Checks Before AI Enters Clinical Reasoning — moved up from the print sidebar
Before AI becomes part of clinical reasoning, a practice should be able to answer six questions.
- Use — What exact task is the system performing: summarization, evidence retrieval, differential support, test selection or treatment guidance?
- Data — What patient information enters the system, where is it stored, who can access it and how is it retained or reused?
- Independence — Can the clinician and AI form meaningfully separate assessments before either sees the other’s conclusion?
- Evidence — Can the practitioner inspect the sources, assumptions, missing information, contradictory findings and uncertainty behind the output?
- Adjudication — Who has the authority and qualifications to accept, modify or reject the recommendation and document the final reasoning?
- Monitoring — How will the practice track errors, overrides, privacy incidents, product updates and patient outcomes?
Practice Protected
FDA’s January 2026 final guidance explains when certain clinical decision-support functions may fall outside the federal definition of a medical device. Enabling a healthcare professional to independently review the basis for a recommendation is one of the relevant statutory criteria. The guidance describes FDA’s current thinking and is nonbinding. (FDA)
In July, FDA requested input on patient safety, education, competency and best practices for certain non-device software functions. The posted comment period closed on August 13. As of August 14, FDA’s page did not yet list a 2026 report. (FDA)
Non-device status does not mean FDA approved. It does not establish that a product is clinically effective, secure or suitable for every use.
For HIPAA-covered entities and business associates, a cloud provider that creates, receives, maintains or transmits ePHI on their behalf generally requires an appropriate business associate agreement. The practice must still complete its own risk analysis. (HHS Office for Civil Rights)
The tools will keep changing. Model rankings will change with them. The harder—and more durable—advantage is the reasoning discipline built inside the practitioner and the practice.
That discipline means thinking independently, using another intelligence to challenge that thinking and still recognizing when the evidence is not good enough to act.
AI is not making clinical reasoning less important. It is raising the standard for it.
Dr. Z’s Take
What interests me most about this research isn’t whether AI can beat a doctor on a test.
What happens when the doctor starts trusting the answer more than the reasoning?
I’ve spent years teaching practitioners how to think through complex cases. Functional medicine is full of patients who don’t hand you one neat diagnosis. They bring years of symptoms, labs, medications, failed treatments, environmental exposures, conflicting opinions and a story that rarely fits into one box.
AI can be incredibly useful in that environment. It can hold an extraordinary amount of information at once. It can find patterns we may have missed, remind us of something that doesn’t fit and challenge a theory we’ve become too attached to.
I believe in that potential so strongly that we invested more than $1 million building a clinical case-planning tool for our own practitioners. We have five software engineers and ten clinical team members working on it.
And what we learned in the process changed the way I think about AI in clinical care.
We couldn’t simply open the consumer version of ChatGPT, Claude or another general-purpose AI tool, feed it protected patient information and ask it to build a case plan. Practitioners have very real HIPAA obligations around where protected health information goes and how it is handled. Putting identifiable patient information into a tool without the appropriate HIPAA protections can create serious privacy and compliance exposure.
But privacy wasn’t the only problem.
Generative AI hallucinates. It can misunderstand information. It can invent facts, sources or connections. And perhaps most dangerously, it can be wrong and sound remarkably confident while doing it.
That’s an inconvenience when you’re asking AI to write an email.
It’s an entirely different problem when you’re making decisions about someone’s health.
So we built something different. AI assists with parts of the process, but we deliberately retained the clinical decision-making. The goal was never to replace the practitioner. It was to make the practitioner dramatically better at doing what humans are already supposed to do.
I sometimes describe it as creating a super practitioner.
Imagine being able to look across a complicated patient history without forgetting the laboratory result from three years ago, the medication that changed six months later, the symptom that appeared after an infection or the intervention that helped one system while making another worse. Technology can help organize all of that complexity, while the clinician still determines what matters, what fits, what doesn’t and what should happen next.
That is where I think AI becomes incredibly powerful.
But it can also make a weak theory sound brilliant. If I enter a case already convinced the patient has mold illness, Lyme, MCAS or mitochondrial dysfunction and then ask AI to support that idea, I may get the most beautifully organized confirmation bias I’ve ever seen.
That isn’t better medicine.
For me, the real value of AI is not having another voice agree with me. It’s having another intelligence help challenge me.
What am I missing? What doesn’t fit? What evidence argues against my theory? What information would change my mind? What should I not do yet?
Those are the questions I want AI helping us answer.
Maybe someday the technology reaches a point where I’m comfortable loosening the grip much further. We are not there today. Not based on the evidence I’m seeing, and certainly not based on what we’ve learned building a system designed to support complex clinical reasoning.
AI doesn’t make the practitioner less important. I think it makes excellent clinical judgment more important. Information is becoming almost effortless to generate. Knowing which information deserves to influence the care of the human being sitting in front of you is something else entirely.
The practitioner of the future won’t win by memorizing more than the machine.
They’ll win by knowing what to ask, what to question, what to ignore—and when the smartest answer is still, “We don’t know enough yet.”
Frequently Asked Questions
Is AI better than doctors at diagnosis?
On some controlled tests, yes. A system evaluated on a benchmark built from New England Journal of Medicine clinicopathological conference cases reported 80% diagnostic accuracy against a 20% average among participating generalist physicians, and in a separate trial a standalone model scored 16 percentage points higher than physicians using conventional resources. Both were simulated conditions. Neither showed better bedside care or better patient outcomes.
Does giving a doctor access to AI improve their diagnostic accuracy?
Not on its own. In a randomized trial of 50 physicians working through difficult vignettes, median diagnostic-reasoning scores were 76% with GPT-4 and 74% with conventional resources — not a statistically significant difference. Capability does not transfer automatically from a tool to the person using it.
What made the difference in the trials where performance improved?
Both changed more than access. Physicians who first completed a 20-hour AI-literacy curriculum scored 71.4% with LLM access against 42.6% without. A trial using a purpose-built workflow that gathered clinician and AI assessments separately, compared them and critiqued both scored 85% against 75% for conventional resources. Neither result transfers to an ordinary commercial chatbot.
Can AI make a clinician’s reasoning worse?
Yes. Systematically biased AI predictions reduced clinician diagnostic accuracy by 11.3 percentage points in a randomized vignette study, and adding image-based explanations did not significantly remove the damage. Even in the workflow trial that raised average performance, actionable scores fell in 12 cases.
Does an AI second opinion stay independent?
Only if it is structured to. In a post hoc analysis of 58 matched cases, complete overlap between the clinician’s initial diagnoses and the AI output was 3% when the AI worked first, and 48% when the AI saw the clinician’s assessment first. A second opinion loses much of its value once one side knows what the other thinks.
Can I put patient information into ChatGPT or Claude?
Practitioners have HIPAA obligations governing where protected health information goes and how it is handled. Putting identifiable patient information into a general-purpose tool without the appropriate protections can create serious privacy and compliance exposure. For covered entities, a cloud provider handling ePHI on their behalf generally requires a business associate agreement, and the practice must still complete its own risk analysis.
Is AI clinical decision support regulated as a medical device?
Sometimes. FDA’s January 2026 final guidance explains when certain clinical decision-support functions may fall outside the federal definition of a medical device; whether a healthcare professional can independently review the basis for a recommendation is one of the relevant statutory criteria. The guidance is nonbinding. Non-device status does not mean FDA approved, and it establishes nothing about whether a product is effective, secure or suitable for a given use.
What should a practice check before using AI in clinical reasoning?
Six things: the exact task the system performs; what patient data enters it and how it is stored and reused; whether the clinician and AI can form separate assessments before either sees the other; whether the sources, assumptions and uncertainty behind the output can be inspected; who is authorized to accept, modify or reject a recommendation; and how the practice will monitor errors, overrides, privacy incidents and outcomes.
Is there evidence AI improves outcomes in functional medicine or chronic illness?
No. The evidence reviewed for this article included no trial validating an AI-supported reasoning system specifically for functional medicine or complex chronic illness. The closest pragmatic evidence — 16 Kenyan primary-care facilities, 103 clinical officers, 9,691 patients — improved documentation quality without changing 14-day treatment failure.
Subscribe
Subscribe for free to get the latest insights in your inbox every week.