Guides & How-Tos

How we think about clinical AI research

2026-07-20 · 6 min read

There’s been a lot of research published on clinical AI this summer, and it’s worth talking about.

In June, a Nature Medicine study compared general-purpose language models against dedicated clinical AI tools — OpenEvidence, UpToDate Expert AI — and found that the general-purpose models outperformed them across medical knowledge benchmarks and real physician queries. Weeks later, OpenEvidence published its own evaluation showing the opposite result: its specialized tool winning across all five dimensions when 149 physicians graded responses to 620 real point-of-care questions. The Stanford-Harvard NOHARM benchmark ranked AMBOSS first in clinical safety. Doximity claimed the top spot on a newer version of that same benchmark. And today, OpenEvidence cited the NOHARM data as evidence that physicians reach for it more than every other AI chatbot combined.

Each of these companies put out a press release. Each declared some version of victory. The interesting thing is that none of them are necessarily wrong. They’re each measuring something different, in a context that suits their own architecture, and finding what you’d expect them to find.

What the studies are actually measuring

This is worth unpacking briefly, because the details matter more than the headlines.

MedQA tests knowledge recall using licensing exam questions. HealthBench scores multi-turn conversations with clinician evaluators across 60 countries. NOHARM uses real primary care consultation cases and evaluates whether AI recommendations could harm a patient. OpenEvidence’s Real-POCQi benchmark draws from questions physicians actually submitted to its platform.

Each captures something real. But a tool can perform well on one and poorly on another, because the question changed. When a company builds a benchmark from queries submitted to its own platform, the questions are genuine and the physicians grading them are practicing clinicians. But the evaluation is also measuring performance in the environment most familiar to that company’s architecture. It’s just something to keep in mind when reading the results.

The Nature Medicine authors were transparent about a related issue: the frontier models they tested may have seen MedQA and HealthBench questions during training. Their primary evidence came from a third benchmark, built from de-identified clinician queries, to avoid that contamination. Even so, those queries came from one institution in one country.

The context question

Here’s something that’s been on our mind as we’ve followed this research. Every benchmark we’ve come across assumes, at some level, a single correct output. Clinical practice rarely works that way.

A question about first-line antidepressant selection will have a different right answer depending on whether the physician has access to a psychiatrist within two weeks or within six months. Drug coverage varies by province, by plan, by formulary. Guidelines differ between countries, and even where they agree on the evidence, the way they’re applied in a community without subspecialty backup is not the same as the way they’re applied at a teaching hospital.

When a study reports that Tool A outperformed Tool B on 500 MedQA questions, it tells you something useful about knowledge retrieval. It doesn’t say much about whether that tool will help a family physician in rural New Brunswick manage a complex patient with the resources she actually has.

We don’t think this is a flaw in the research. It’s a structural feature of clinical work that benchmarks weren’t designed to capture. Context — the clinical environment, the guideline landscape, the resource constraints a physician is operating within — is the thing that makes a response useful rather than just accurate.

How we are approaching this question

Asklepius is built for Canadian clincian. That’s a deliberate design choice, and it shapes everything from data source layers to how we evaluate quality. Our responses are grounded in Canadian clinical practice guidelines and CPS and Health Canada drug monographs. When a physician asks about a medication, the answer reflects Canadian formulary data rather than defaulting to U.S. sources.

That grounding doesn’t produce a higher number on a U.S. medical licensing exam. It produces responses that are more relevant in the specific context where our clinicians practice.

On the validation side, we run blinded comparative studies where practicing physicians evaluate Asklepius responses alongside other tools. We run ongoing specialist annotation — clinicians reviewing outputs and flagging where the tool fell short, cited the wrong guideline, or missed a clinically important nuance. And we run safety reviews looking specifically for dangerous or misleading responses.

The purpose of that work is to find failures. When a specialist annotator flags that a response missed a drug interaction or defaulted to an American guideline where a Canadian one existed and differed, that goes directly back into the system. It’s continuous, not point-in-time.

We’ve written before about escalation intelligence — teaching Asklepius to recognize when a question exceeds what AI should be answering on its own, and suggesting the physician consult a specialist through Virtual Hallway’s network. That’s not the kind of thing that shows up on a leaderboard. But we think it might matter more than most of the things that do.

Reading the research

We’re not going to tell anyone which tool to use. Clinicians are sharp evaluators of their own tools, and the right answer depends on where and how you practice.

But the next time a clinical AI study lands, a few questions are worth keeping in mind: Were the test questions drawn from real practice or from exam banks? Did the evaluation reflect the clinical context you work in — your country, your guidelines, your patient population, your available resources? Was the study designed independently of the companies being evaluated? And beyond the score itself: does this company have a process for catching the errors that benchmarks miss — not just at publication, but continuously?

A benchmark tells you what a tool was optimized for. That’s worth knowing. What happens after — whether there’s a clinical team reviewing outputs, a feedback loop from real physicians, a mechanism for the tool to keep getting better in the context that matters to you — is harder to measure and, we think, at least as important.

We’d rather build something that keeps improving for the physicians who use it than something that wins a comparison designed somewhere else. The research is how we get there.


← Back to Resources