📑 Table of contents

"Humans in the loop degrade AI performance": Emanuel and Khosla flip the medical debate in JAMA

Actu IA 🟢 Beginner ⏱️ 12 min read 📅 2026-08-29

"Humans in the loop degrade AI performance": Emanuel and Khosla flip the medical debate in JAMA

🔎 A converted skeptic publishes the year's most controversial health article

Ezekiel Emanuel is no Silicon Valley fanboy. An architect of Obamacare and a respected ethicist at the University of Penn, he had previously dismissed medical AI as "bullshit." In August 2026, he co-authors an op-ed in JAMA with Vinod Khosla arguing that autonomous AI will outperform doctors — and that humans in the loop degrade performance.

The timing is not coincidental. Models like GPT-5.5 (OpenAI) are hitting 98.2 on agentic benchmarks, and Anthropic's Claude Opus 4.7 (Adaptive) surpasses 94. Reasoning capabilities have exploded in eighteen months.

What changes everything is the method. Emanuel and Khosla aren't inventing anything: they review all publications on AI in medicine since January 2024. Their conclusion is unequivocal across five fundamental clinical tasks.


The Essentials

  • The August 2026 JAMA op-ed (Emanuel, Baker-Butler, N. Khosla, V. Khosla) asserts that autonomous AI will outperform physician+AI systems on five tasks: information gathering, test ordering, results interpretation, diagnosis formulation, and treatment plan development.
  • The shocking statement: "humans in the loop degrade AI performance" — human intervention reduces AI performance rather than improving it.
  • Ezekiel Emanuel, a former critic of medical AI, has been swayed by the data and by his colleague Robert Wachter (UCSF), who himself converted after direct clinical experiences.
  • The AMA and Dr. John Whyte strongly push back, pointing out the lack of real clinical trials and a February 2026 Oxford study showing that patients do not know how to converse effectively with LLMs.
  • The debate pits two visions against each other: AI as an augmentation tool vs. AI as an autonomous replacement.

The 5 tasks where autonomous AI dominates, according to JAMA

The argument of the op-ed is based on a precise breakdown of medical work into five distinct cognitive tasks. For each one, the authors cite publications showing that AI alone outperforms the physician+AI duo.

1. Gathering patient information

A medical LLM asks the right questions exhaustively, without fatigue, without mood biases, and without forgetting to ask about a symptom because it is rushed by the next patient.

Current generalist models like Gemini 3.1 Pro (Google), which scores 92 on the general benchmark, are capable of conducting a structured clinical interview in a few minutes.

2. Ordering additional tests

Fewer unnecessary tests, greater relevance. AI does not prescribe a scan to reassure the patient or cover itself legally. It prescribes what is statistically relevant.

3. Interpreting results

Imaging, biology, genomics — AI reads and cross-references data simultaneously, whereas a human processes them sequentially and might miss an interaction between two seemingly normal results.

4. Formulating the diagnosis

This is the core of the profession. And this is where the gap is the clearest according to the compiled data. A model like GPT-5.5, with its agentic score of 98.2, identifies patterns that physicians miss, including in rare cases.

5. Developing the treatment plan

Integration of the latest guidelines, the patient's contraindications, drug interactions, preferences — all in a single pass.


Why "humans in the loop" degrades performance

The most cited phrase from the article — "humans in the loop degrade AI performance" — relies on a specific mechanism documented in several studies.

When a doctor verifies the AI's output, two things happen. First, they introduce their own cognitive biases: anchoring, confirmation, overconfidence. Second, they modify the AI's recommendations even when they were correct.

Stanford Medicine had already documented this phenomenon in February 2025: chatbots alone outperformed doctors on nuanced clinical decisions, but doctor+AI teams showed mixed results.

The explanation is psychological. The doctor doesn't just verify — they reinterpret. And this reinterpretation, even when done by an expert, introduces noise into a system that had already found the right answer.

Emanuel and Khosla argue that forcing a human into the loop, as regulators demand, amounts to imposing an unnecessary cognitive tax. The Decoder summarizes this position: regulators should not mandate a human in the loop if the data shows that it degrades the outcome.


Ezekiel Emanuel's reversal: from "bullshit" to JAMA

This is perhaps the most striking detail of the whole affair. Emanuel isn't a venture capitalist desperate for buzz. He is a doctor, bioethicist, and former special advisor to Obama.

In his interview with Chief Healthcare Executive, he admits it bluntly: he used to call medical AI "bullshit." What changed his mind wasn't Google Keynotes or OpenAI demos. It was Robert Wachter.

Wachter, a professor of medicine at UCSF, was also a skeptic. He changed his mind after testing the models in real clinical conditions and convinced Emanuel to look at the data without ideological filter.

Emanuel now believes that "we are going to have autonomous clinical AI." The phrasing is remarkable: not "we should," but "we are going to." He isn't making a normative prescription — he is describing what he sees happening.


Khosla : « Game mostly over for human doctors vs AI »

Vinod Khosla, for his part, has never been one to hold back. On X, he summarizes the op-ed in a single sentence: « Game mostly over for human doctors vs AI. »

This is consistent with his historical position. Khosla has argued since 2012 that 80% of medical work will be replaced by machines. The difference in 2026 is that he is no longer alone in saying it, and he has a co-author who lends academic credibility to the thesis.

On the Khosla Ventures website, the op-ed is published in full alongside the counterpoint from the ACP (American College of Physicians). The ACP argues that AI « should be limited to a support role in clinical decision-making » — exactly what the op-ed contests.

WSJ Pro Venture Capital notes that the article argues AI rivals or outperforms doctors on these five cognitive tasks. Not « might one day », but « rivals or outperforms » in the present tense.


The models behind the argument: what changed in 2025-2026

JAMA's argument would not exist without the explosive progress of models. The numbers speak for themselves.

Model Publisher Agentic Score General Score
GPT-5.5 OpenAI 98.2 91
Claude Opus 4.7 (Adaptive) Anthropic 94.3 90
Gemini 3 Pro Deep Think Google 95.4 90
Claude Opus 4.6 Anthropic 84.7 87
DeepSeek V4 Pro (Max) DeepSeek 88

These scores are not chat benchmarks. They are measures of the ability to execute complex reasoning chains, to plan, to correct its own errors — exactly the type of cognition required for the five identified medical tasks.

Claude Opus 4.7, with its adaptive mode, is particularly relevant in the medical context: it adjusts its reasoning level to the complexity of the case. A simple case receives a quick response; an atypical case triggers deep reasoning.

This evolution explains why Digital Health Wire can write that AI is "already better" than doctors at the identified tasks. The benchmarks are no longer theoretical — they correspond to operational capabilities.

What makes this debate possible is also the work on alignment. Anthropic, for example, unveiled an automated alignment researcher that outperforms humans on benchmarks, fixing 10/10 alignment issues at a cost 40x lower. The reliability of outputs is no longer the bottleneck.


The controversy: what critics reproach the op-ed for

The JAMA article immediately drew mixed reactions. The praise focuses on the rigor of the literature review. The criticisms, on the other hand, are structured and deserve to be examined.

The problem of simulations vs. clinical reality

This is the main objection from the AMA and Dr. John Whyte. The studies compiled by Emanuel and Khosla are mostly retrospective experiments or simulations. None is a real-world randomized clinical trial.

A diagnosis on a paper case is not a diagnosis at the bedside of an anxious patient who cannot clearly describe their pain. Fierce Healthcare notes that the debate is growing over the future role of doctors precisely because the gap between simulation and reality remains immense.

The Oxford study from February 2026: patients don't know how to talk to LLMs

The Oxford study is the strongest counter-argument. It highlights three concrete problems: incorrect diagnoses, an inability to understand the urgency of a situation, and above all — patients who do not know how to converse with LLMs.

This is a point that the JAMA op-ed underestimates. AI can be perfect on paper; if the patient cannot communicate their symptoms to it correctly, everything falls apart. Information gathering (task #1) requires an interlocutor capable of expressing themselves clearly.

Khosla's conflict of interest

Vinod Khosla is one of the biggest investors in healthtech. He has funded dozens of medical AI startups. Digg highlights the skeptical reactions regarding this article co-signed by someone who has a direct financial interest in the thesis they are defending.

Emanuel, for his part, does not have this conflict. This is probably what gives the op-ed its credibility beyond the venture capital sphere.


Wired and the existential question: "What's left for us?"

Wired quotes doctors asking an existential question in the face of these results. If AI performs better on all five cognitive tasks, what is left for the doctor?

The answer emerging from the debate is not in the JAMA op-ed itself, but in the reactions it provokes. What remains, according to defenders of the human role, is:

  • The therapeutic relationship. A patient does not want to be diagnosed by a machine; they want to be cared for by a human who looks them in the eye.
  • Navigating radical uncertainty. Cases where data is insufficient, where clinical intuition — that unmeasurable thing — makes the difference.
  • Case-by-case ethics. Guidelines are averages. Patients are individuals.

The problem is that these arguments are difficult to quantify. And Emanuel, a data man, blames them for it.

As NBTX News notes in its coverage, the question is no longer theoretical. American doctors are already changing specialties, anticipating a future where diagnosis and treatment planning will no longer be their job.

This movement is not isolated. In other sectors, Allianz eliminated 1,800 jobs because of AI, the most concrete signal of the massive replacement looming. The parallel with medicine is inevitable.


Medical AI is not just about diagnosis: the Cambridge vaccine

The JAMA debate focuses on diagnosis and treatment. But medical AI is advancing on other fronts where the question of the "human in the loop" arises differently.

The first universal vaccine designed entirely by AI enters clinical trials in Cambridge, with results confirming safety. There, AI did not replace a doctor — it replaced years of research in molecular biology.

These two dynamics (cognitive replacement and research augmentation) reinforce each other. The AI that diagnoses better will be fed by the treatments that AI itself helped to discover.


❌ Common mistakes in reading this debate

Mistake 1: Confusing "AI does better" with "AI does everything"

The JAMA op-ed focuses on five specific cognitive tasks. It does not say that AI can replace a surgeon in the operating room, nor that it handles life-threatening emergencies with the same efficiency. Reducing the debate to "AI replaces doctors" is a dangerous oversimplification from both sides.

Mistake 2: Ignoring Khosla's conflict of interest

Citing the op-ed without mentioning that Vinod Khosla is a major healthcare AI investor means presenting an argument from authority without context. The credibility comes from Emanuel, not Khosla. Separating them intellectually is necessary.

Mistake 3: Taking benchmark scores for real clinical performance

A score of 98.2 on the GPT-5.5 agentic benchmark is impressive. But it is not a correct diagnosis rate in real-world conditions. Benchmarks measure reasoning ability in standardized formats — not the management of a diabetic, depressed, and non-compliant patient.

Mistake 4: Binarily opposing AI vs. doctors

The Stanford study from February 2025 already showed that reality is nuanced: sometimes AI alone wins, sometimes the doctor alone wins, sometimes the duo wins, sometimes the duo loses. The "human in the loop = always better" model is false. But the "human in the loop = always worse" model is too.

❓ Frequently Asked Questions

Who are the authors of the JAMA op-ed?

Ezekiel Emanuel (University of Penn), Baker-Butler, Nayana Khosla, and Vinod Khosla. The article was published on August 17, 2026, in JAMA, with a counterpoint from the American College of Physicians.

What exactly does "humans in the loop degrade AI performance" mean?

That when a doctor verifies or modifies the output of an AI system, the final result is statistically worse than the unmodified AI output. The main mechanism is the introduction of human cognitive biases into a system that had found the right answer.

Does the AMA support the article's thesis?

No. The American College of Physicians, in its counterpoint published alongside the op-ed, asserts that AI must remain limited to a support role in clinical decision-making. The AMA has not yet taken an official stance on the clinical autonomy of AI.

What are the main limitations of the argument?

The absence of randomized clinical trials in real-world conditions, Vinod Khosla's conflict of interest, and the Oxford study showing that patients do not know how to interact effectively with medical chatbots.

Are there cases where doctor+AI is better than AI alone?

Yes. The Stanford study from February 2025 showed mixed results for doctor+AI teams. The problem is that we do not yet know how to predict in which cases human input is beneficial and in which cases it is detrimental.


✅ Conclusion

The op-ed by Emanuel and Khosla in JAMA does not mark the end of doctors — it marks the end of the dogma that a human in the loop always improves AI systems. Across five medical cognitive tasks, the data accumulated since 2024 shows the opposite. What remains to be seen is whether these data hold up in the transition from simulation to clinical reality, which is where the Oxford study and the AMA's criticisms find their strength. The debate is no longer theoretical: it is temporal.