Guides & How-Tos

Evidence was never the hardest problem

2026-07-13 · 5 min read
AI is best when it can escalate to humans

If you follow clinical AI right now, you’ve seen the race. Every few weeks a new comparison: which model retrieves guidelines faster, which one synthesizes the literature better, which one cites more sources. The benchmarks keep improving, and honestly, the results are impressive. Ask a well-formed question and within seconds you have prescribing information, guideline recommendations, and primary literature summarized better than most of us could manage on our own.

We’ve spent a lot of effort making Asklepius good at exactly this. Sourced responses, Canadian guidelines, CPS monograph data at the point of query. We think it matters, and our validation work suggests physicians agree.

But the longer we build, the more we’ve come to believe that access to evidence was never the hardest problem in clinical work. It may not even be the real problem.

The question after the evidence

Every clinician knows this moment. The guideline is open in front of you. You know what the next-line treatment is. And something still isn’t right.

The patient is older than anyone enrolled in the trial. They have three conditions the guideline barely mentions. Their kidney function complicates the first-line option, their family situation changes what’s realistic, and they have preferences of their own. The evidence hasn’t become less valuable — it just isn’t the whole story. Evidence narrows the possibilities. It informs the decision. It doesn’t make it.

Medical school teaches us evidence: landmark trials, confidence intervals, systematic reviews. Then the first patient walks through the door and we start learning the other thing — the interpretive work of deciding whether the evidence applies to this patient, in this situation, right now. That work is judgment, and it’s built from experience, pattern recognition, and knowledge that doesn’t live in any database.

Two kinds of clinical questions

Once you see this distinction, you notice that clinical questions come in two kinds.

Some are evidence questions. What does the Canadian guideline recommend for first-line therapy? Is there a clinically significant interaction between these two drugs? What’s the cross-taper protocol? These questions have answers that live in maintained references, and a well-built AI tool can surface them quickly, with sources, so you can verify them yourself. This is what Asklepius is for.

Others are judgment questions. Does this recommendation apply to my 84-year-old with an eGFR of 30 and a strong opinion about pill burden? Should this atypical presentation change my plan, or is it noise? When does my uncertainty warrant acting differently rather than simply being acknowledged? These questions aren’t answered by retrieving more evidence. They’re answered by clinical reasoning — often by the accumulated experience of someone who has seen this exact ambiguity two hundred times.

The problem is that so far, the field has tried to treat both kinds of questions identically. Same interface, same confident tone, same fluent paragraph — whether the question is a textbook lookup or a genuine judgment call sitting at the edge of the evidence. The model has no way to tell you which kind of question you just asked, and no way to act differently even if it could.

That failure mode worries us more than an occasional wrong citation. A wrong citation can be checked. A judgment question dressed up as a settled answer is harder to catch, because nothing about the response signals that you’ve left the territory where a database is the right tool.

Why we’re building an escalation layer

This is the design problem behind what we call escalation intelligence, which we first wrote about when we introduced Asklepius. The idea is simple to state and hard to build: train the system to recognize when a question is drifting from evidence into judgment — and when it does, don’t generate a more elaborate paragraph. Offer a colleague.

This is a little bit easier for us to build because Asklepius is built on Virtual Hallway, a network of more than 10,000 clinicians with specialists across every major discipline. When a question looks like it warrants specialist input, the next step isn’t a suggestion to “consider consulting cardiology” that leads nowhere. It’s a consultation pathway that already exists, with a specialist who can engage with the parts of the question no reference tool can: the atypical presentation, the competing priorities, the uncertainty that should change the plan.

Most AI tools have no way to act on the realization that a question needs a human. There’s nowhere for the question to go, so the model gives its best answer regardless. We think the more honest architecture is one where handing off is a first-class outcome — where the system’s most valuable output, for certain questions, is a connection rather than a completion.

Where this stands, honestly

Escalation intelligence is still in an early phase, and we want to be straightforward about why it’s hard: recognizing the boundary between an evidence question and a judgment question is itself a judgment problem. There’s no clean rule. A medication dose can be a lookup for one patient and a genuine dilemma for another.

So we’re approaching it the way we’ve approached validation from the start — with specialists in the loop. We’re running annotation studies where physicians across specialties review real (de-identified) clinical questions and tell us which ones they’d want a colleague on. That specialist-labelled data is how the system learns what “this warrants a conversation” looks like across disciplines. We’d rather the system escalate too readily than too rarely, and we’ll publish what we learn as the work matures.

What doesn’t change

None of this moves the judgment out of your hands. Asklepius is an educational and informational tool for qualified healthcare professionals. It doesn’t make clinical decisions, and it doesn’t replace clinician-to-clinician consultation, specialist referral, or your own clinical judgment — the escalation layer exists precisely because those things can’t be replaced. When Asklepius surfaces evidence, you verify it and decide whether it applies. When it suggests a question may warrant a specialist, you decide whether to make the call.


← Back to Resources