Toward the Autonomous AI Doctor: What the 500-Patient Doctronic Study Really Shows

Can an artificial intelligence system actually function like a doctor?

A 2025 preprint from the Doctronic Research Group offers one of the more interesting real-world tests of that question. Researchers evaluated a proprietary multi-agent artificial intelligence system, Doctronic, across 500 consecutive urgent-care telehealth encounters and compared its diagnostic assessments and treatment plans with those produced by board-certified clinicians.



The headline numbers are striking: the AI's top diagnosis matched the clinician's diagnosis in 81% of cases, at least one of its top four diagnoses matched in 95.4% of cases, and the AI and clinician treatment plans were considered clinically compatible in 99.2% of encounters.

But there is an important distinction between agreement and accuracy.

The researchers themselves acknowledge that the study was designed primarily to measure concordance, not whether either the AI or the human clinician was ultimately correct based on a definitive diagnosis or subsequent patient outcome.

Bottom line: The study provides an important signal that a carefully engineered multi-agent AI system can produce clinical assessments broadly comparable with those of clinicians in a limited urgent-care telehealth setting. It does not establish that AI doctors are equivalent to human doctors, nor that a 99.2% treatment-plan concordance rate represents 99.2% clinical accuracy or safety.

What Was the Doctronic Study?

The study, titled Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting, was posted on medRxiv on July 16, 2025.

The authors analyzed 500 fully de-identified, nonemergency urgent-care telehealth encounters from the first week of March 2025.

Each encounter involved two evaluations:

  1. The patient first interacted with Doctronic, operating autonomously.
  2. The patient was subsequently evaluated by a board-certified clinician.

The researchers then compared the AI-generated and clinician-generated clinical documentation, diagnoses and management plans.

Importantly, clinicians were permitted, although not required, to use the AI-generated documentation when conducting their assessment.

Read the full medRxiv preprint

What Is Doctronic?

Doctronic is described by the researchers as a cloud-based, modular multi-agent AI system incorporating more than 100 LLM-powered agents.

Rather than relying on a single large language model to produce an answer, the architecture divides clinical functions among multiple agents with different roles.

The system is designed to conduct a medical history, summarize the clinical information, generate a differential diagnosis and produce a treatment or diagnostic plan.

The output is structured into a SOAP-style clinical note.

This distinction matters because the study is not simply testing a patient asking a general-purpose chatbot, "What do my symptoms mean?" It is evaluating an engineered workflow intended to mimic several components of a clinical care team.

The Main Results

Measure Study result What it means
Top-1 diagnostic concordance 81% (405/500) The AI's primary diagnosis matched the clinician's primary diagnosis.
Top-4 diagnostic concordance 95.4% (477/500) At least one diagnosis in the AI's top four matched a clinician diagnosis.
Treatment-plan concordance 99.2% (496/500) The plans were judged clinically compatible and guideline-concordant by the study's blinded LLM judge.
Cases with substantive treatment-plan divergence 0.8% (4/500) The researchers reported no immediate or high-harm risk in these divergent cases.
Reported clinical hallucinations 0 cases No fabricated diagnosis or unsupported treatment was identified under the study's definition.

These numbers come directly from the preprint. The 81% result corresponds to 405 of 500 encounters, while the 95.4% result corresponds to 477 of 500. The 99.2% treatment-plan result represents 496 of 500 encounters.

81% Diagnostic Agreement Is Not the Same as 81% Accuracy

This is one of the most important points for interpreting the study.

The study did not independently establish a definitive "ground truth" diagnosis for every patient and then determine how often AI and doctors were correct.

Instead, the primary comparison asked whether the AI and clinician reached the same or clinically compatible conclusions.

That means the study can demonstrate concordance, but it cannot by itself tell us whether both parties were correct.

For example, if an AI system and a physician both independently make the same incorrect diagnosis, the pair would score as concordant even though the clinical decision would not be correct.

Key distinction: Agreement with doctors is useful, but agreement with doctors is not a substitute for validation against clinical outcomes or an independently established reference standard.

Why Was Treatment-Plan Concordance 99.2%?

The 99.2% figure is arguably the most attention-grabbing result in the study.

However, the wording needs to be precise.

The researchers did not report that the AI prescribed exactly the same treatment as a doctor in 99.2% of cases.

Instead, treatment plans were considered clinically compatible when the evaluator judged that differences such as equivalent medications, synonyms or different but compatible management approaches would likely lead to similar therapeutic outcomes.

In four of the 500 cases, the treatment plans were classified as substantively divergent.

The study reported that none of those divergences created immediate or high-harm risk.

Therefore, the appropriate description is:

"99.2% treatment-plan compatibility in this study."

That is more accurate than saying:

"The AI was 99.2% as accurate as doctors."

The Role of the LLM Judge

One of the study's most important methodological features is also one of its biggest limitations.

The investigators used a large language model to evaluate many of the comparisons between AI and human clinical notes. The judging system used GPT-4.0 to assess diagnostic matching and treatment-plan compatibility.

The researchers also included human review of discordant cases as a safeguard.

However, this creates an unusual evaluation structure:

AI system → compared with physician → evaluated in part by another AI system.

That is not the same thing as having an independent panel of specialist physicians determine the correct diagnosis for every case using subsequent patient outcomes, imaging, pathology, laboratory confirmation or long-term follow-up.

LLM-based evaluation may be useful for large-scale benchmarking, but it can introduce its own interpretation errors or biases. The authors explicitly acknowledge this limitation in the manuscript.

What Happened in the 97 Discordant Cases?

The investigators identified 97 encounters in which the AI and clinician were not judged to have the same primary diagnosis.

A board-certified physician then manually reviewed these discordant pairs.

According to the study:

  • In 35 cases (36.1%), the AI-generated reasoning was judged superior.
  • In 9 cases (9.3%), the human clinician was judged superior.
  • In 36 cases (37.1%), the diagnoses were actually considered clinically similar despite different wording or specificity.

The remaining cases involved ambiguity or other differences in interpretation.

This result is interesting, but it should not be translated into a claim that "AI doctors beat doctors 4-to-1."

The comparison was conducted specifically among cases already identified as diagnostically discordant, and the study's own methodology and definitions shaped the classification.

The Biggest Methodological Concern: Anchoring Bias

Perhaps the most important limitation is that the human clinician could see the AI-generated documentation before completing the clinical encounter.

This creates the possibility of anchoring bias.

An AI-generated diagnosis can influence what a physician asks about, what possibilities the physician considers and how the physician interprets subsequent information.

Even if clinicians were not required to accept the AI assessment, exposure to it could potentially increase agreement between the two outputs.

This is particularly important when interpreting the 99.2% concordance result.

The authors themselves acknowledge this concern and recommend future testing in which clinicians assess patients without first seeing the AI-generated note.

Another Problem: No Patient-Outcome Gold Standard

The researchers explicitly state that the investigation was based on concordance rather than correctness.

The study did not establish whether the AI's treatment recommendations actually produced better patient outcomes over time.

A stronger clinical validation study would ideally incorporate measures such as:

  • confirmed final diagnoses;
  • follow-up outcomes;
  • hospital admissions or emergency visits after the encounter;
  • medication-related adverse events;
  • missed diagnoses;
  • unnecessary investigations;
  • patient recovery and symptom resolution; and
  • longer-term safety outcomes.

Those endpoints would move the evidence from "AI agreed with clinicians" toward the more important question: "Did the AI make the right clinical decision for the patient?"

How Representative Were the Patients?

The study evaluated 500 U.S. urgent-care telehealth encounters and reported presentations spanning more than 100 major ICD-10 diagnostic categories.

The authors describe the sample as broadly distributed across age, sex, presenting complaints and comorbidity complexity.

However, the findings should not automatically be generalized to every medical setting.

A text-based urgent-care telehealth encounter is fundamentally different from:

  • emergency medicine;
  • intensive care;
  • oncology;
  • surgery;
  • obstetrics;
  • pediatrics;
  • psychiatry;
  • critical care;
  • hospital medicine; or
  • complex longitudinal primary care.

The authors also acknowledge that the study was conducted in an English-language U.S. telehealth setting, limiting its generalizability to other languages, healthcare systems and clinical environments.

What Does "No Hallucinations" Actually Mean?

The study reported zero clinical hallucinations.

That does not mean the system made zero mistakes.

The authors defined hallucination around fabricated clinical information—for example, a diagnosis or treatment not supported by information in the patient transcript.

A system can avoid inventing information and still:

  • misinterpret a real symptom;
  • miss an important differential diagnosis;
  • underestimate disease severity;
  • prioritize the wrong diagnosis;
  • recommend an unnecessary test;
  • fail to recognize a rare condition; or
  • make a clinically suboptimal decision.

The paper itself reports one technical error in which an AI-generated note lacked sufficient patient data for evaluation.

Therefore, "zero hallucinations" should not be presented as equivalent to "zero clinical errors."

Why the Study Matters Anyway

Despite these limitations, the research is important.

Most discussions of medical AI have focused on models that assist doctors—for example, generating summaries, suggesting differential diagnoses, documenting visits or retrieving information.

Doctronic represents a different direction: an attempt to create an autonomous clinical workflow in which multiple AI agents collectively perform history-taking, reasoning and documentation.

The important scientific question is therefore no longer simply:

"Can an AI answer a medical question?"

It becomes:

"Can an engineered AI system safely perform an end-to-end clinical workflow under realistic conditions?"

This study provides preliminary evidence that the answer may eventually be yes for at least some limited clinical environments.

AI Doctor vs Human Doctor: What This Study Actually Demonstrates

Question What the study supports What it does not prove
Can an AI system generate diagnoses similar to clinicians? Yes, in this dataset. That it is universally as accurate as physicians.
Can an AI generate clinically compatible treatment plans? Yes, in 99.2% of evaluated encounters. That 99.2% of recommendations are independently proven correct.
Can autonomous medical AI avoid hallucinated clinical facts? No hallucinations were identified under the study definition. That the AI has zero risk of clinically important errors.
Can AI replace doctors? The study supports further investigation of autonomous AI. No. The study was not designed to establish physician replacement.

What Would a Stronger Next-Generation Study Look Like?

The next generation of research should address the methodological weaknesses of this preliminary study.

A stronger design would include:

  1. Independent clinician assessment: physicians should evaluate patients without first seeing the AI-generated assessment.
  2. Multiple independent adjudicators: diagnostic correctness should be assessed by appropriately qualified human experts rather than relying primarily on an LLM judge.
  3. Patient outcomes: researchers should determine whether the diagnosis and treatment ultimately produced better clinical outcomes.
  4. Prospective evaluation: patients should be enrolled prospectively rather than relying only on retrospective encounters.
  5. Broader clinical settings: testing should extend beyond urgent-care telehealth.
  6. Adversarial testing: AI should be deliberately tested with rare diseases, ambiguous presentations, conflicting information and misleading symptoms.
  7. Failure analysis: researchers should report clinically significant false negatives, inappropriate reassurance and unnecessary treatment—not just hallucinations.

The Bigger Picture: From AI Assistant to AI Clinician

Medical AI is evolving through several stages.

The first stage was largely information retrieval: finding medical literature, guidelines and facts.

The second stage involved clinical assistance: summarizing notes, documenting visits, suggesting diagnoses and supporting physicians.

The next stage is increasingly focused on agentic clinical systems: AI that can interact with patients, gather information, reason through a differential, generate documentation and recommend management with limited human intervention.

That transition creates a much higher safety requirement.

An inaccurate chatbot answer may be inconvenient. An inaccurate autonomous medical decision can be harmful.

For that reason, the future of AI medicine should not be evaluated solely by benchmark scores or agreement rates.

The central questions should be:

Does it improve patient outcomes?

Does it reduce missed diagnoses?

Does it reduce unnecessary care?

Does it perform safely across different populations?

Can its decisions be audited and explained?

And what happens when the AI is wrong?

Our Assessment

The Doctronic study should be viewed as an important but preliminary milestone in autonomous medical AI.

The 500-case dataset, high treatment-plan concordance and strong diagnostic agreement are encouraging. The study also provides a useful example of how multi-agent architectures might be used to create more structured clinical workflows.

However, the evidence remains limited by several factors:

  • the study is a preprint rather than a peer-reviewed clinical trial;
  • all authors disclosed equity ownership of Doctronic;
  • the main endpoint is concordance rather than clinical correctness;
  • the physician could see the AI-generated documentation, creating potential anchoring bias;
  • an LLM was used to judge important elements of the comparison;
  • there was no patient-outcome validation; and
  • the dataset came from U.S. English-language urgent-care telehealth encounters.

Accordingly, the study does not establish that autonomous AI doctors are ready to replace physicians.

What it does show is arguably more interesting: carefully engineered agentic AI systems may be approaching a level of clinical workflow performance that deserves serious prospective testing.

2026 Update: Where the Evidence Stands

As of September 2026, the source examined here remains identified by medRxiv as a preprint, and I did not find a clearly corresponding peer-reviewed journal publication that supersedes the original version in the searches conducted for this update.

That distinction matters. The study is potentially influential, but readers should continue to treat its numerical findings as preliminary research evidence rather than established clinical evidence.

Evidence interpretation: Promising real-world benchmarking signal, but insufficient by itself to establish autonomous AI doctor safety, diagnostic superiority or equivalence to human physicians.

Frequently Asked Questions

Did the AI doctor outperform human doctors?

Not conclusively. In a subset of 97 diagnostically discordant cases, a physician reviewer rated the AI reasoning superior in 36.1% and the human reasoning superior in 9.3%. However, this was a secondary analysis of discordant cases and should not be interpreted as proof that AI doctors outperform physicians overall.

What does the 99.2% figure mean?

It means that 496 of 500 AI and clinician management plans were judged clinically compatible and guideline-concordant by the study's evaluation process. It does not mean the AI was independently proven 99.2% accurate.

Was the study peer reviewed?

No. The source is a medRxiv preprint and should be regarded as preliminary research until independently peer reviewed and replicated.

Did the study prove that AI can replace doctors?

No. It tested one proprietary AI system in 500 urgent-care telehealth encounters. It did not establish equivalence across specialties, complex disease, emergency care, longitudinal care or real-world patient outcomes.

Why is the study still important?

Because it moves the discussion beyond simple chatbot benchmarks toward testing an autonomous multi-agent clinical workflow in real patient encounters. That is an important direction for future research.

Conclusion

The question is no longer whether artificial intelligence can produce medically plausible answers. Increasingly, researchers are testing whether AI can perform an entire clinical workflow.

The Doctronic study provides encouraging preliminary evidence that an autonomous multi-agent AI system can produce diagnostic assessments and treatment plans that frequently resemble those generated by board-certified clinicians in urgent-care telehealth.

But concordance is not the same as correctness.

The most important next step is not another benchmark showing that AI agrees with doctors. It is prospective, independently adjudicated research demonstrating whether AI-assisted care can actually improve—or, at minimum, safely maintain—patient outcomes.

The future should not be framed as AI versus doctors. The more meaningful goal is AI working alongside doctors: combining the strengths of artificial intelligence with human clinical judgment, experience, empathy and accountability to deliver better, safer and more personalized care.

The question is not whether AI can replace doctors, but whether the best possible combination of AI and doctors can take patient outcomes to the next level—globally, across healthcare systems and across the full spectrum of medical care.

That is the standard that will determine whether the "AI doctor" remains an intriguing technology demonstration or becomes a genuinely useful and transformative component of clinical medicine.


Medical and research disclaimer: This article discusses preliminary medical artificial intelligence research for educational purposes. The cited study is a preprint and has not been established as evidence that autonomous AI systems are safe or appropriate substitutes for qualified healthcare professionals. AI-generated medical information should not be used to diagnose, treat or manage a medical condition without appropriate professional oversight.

Primary Source

Hayat H, Kudrautsau M, Makarov E, et al. Toward the Autonomous AI Doctor: Quantitative Benchmarking of an Autonomous Agentic AI Versus Board-Certified Clinicians in a Real World Setting. medRxiv. 2025. doi:10.1101/2025.07.14.25331406.

View the study DOI | Read the full preprint

Comments

Popular Posts