Definitive category guide

AI evidence-based medicine

AI evidence-based medicine applies computational tools within the established discipline of evidence-based practice: combining the best available research with clinical expertise and each patient’s values, goals and circumstances.

Global, non-jurisdictional guide For qualified clinicians Human judgement remains essential

The foundation

Evidence-based medicine has three inseparable parts

AI can help organise information and make reasoning more explicit. It does not replace the professional integration at the centre of evidence-based care.

Best available evidence

Current, relevant research should be appraised for validity, magnitude, certainty, applicability and limitations—not merely retrieved or summarised.

Clinical expertise

Clinicians interpret incomplete histories, examination findings, trajectories, comorbidities and local constraints that a model may not adequately represent.

Patient values and context

Choices depend on the individual’s preferences, goals, risks, access, culture and circumstances. These cannot be reduced to a generic model output.

Core principle: an answer is not evidence-based simply because it cites a paper. The evidence must be trustworthy, applicable and integrated with clinical judgement and the patient’s priorities.

The three-part foundation follows the foundational description of evidence-based medicine in The BMJ.

Responsible role

Where AI can contribute—and where it cannot

Appropriate support

  • Structure a clinical question and surface information gaps.
  • Generate hypotheses for a clinician to consider and independently verify.
  • Organise relevant findings, alternative explanations and potential red flags.
  • Link claims to inspectable sources when reliable provenance is available.
  • Support explicit probability updates and uncertainty discussions.

Unsafe assumptions

  • That fluent language is equivalent to clinical correctness.
  • That a citation proves the surrounding claim or applies to the patient.
  • That a model can see missing context, recognise every emergency or eliminate bias.
  • That performance in one setting transfers unchanged to another.
  • That automation removes the need for consent, accountability or oversight.

Evaluation framework

Six questions to ask before clinical use

Evaluation should match the intended users, workflow, population and consequences of error. A broad model score is not a substitute for use-case testing.

1. Intended use

Is the user, task, care setting and boundary of the system stated precisely enough to test?

2. Evidence provenance

Can users inspect the sources, dates and transformations behind consequential claims?

3. Clinical performance

Has the complete workflow been evaluated on representative cases using clinically meaningful measures?

4. Uncertainty and calibration

Does the system communicate uncertainty appropriately, and do estimated probabilities match observed outcomes?

5. Safety and equity

Are failure modes, subgroup performance, automation bias, omissions and foreseeable misuse actively examined?

6. Governance

Are responsibility, privacy, access controls, change management, incident response and monitoring defined?

Use the clinical AI evaluation checklist to turn these questions into a structured review.

For a broader governance baseline, see the World Health Organization’s ethics and governance guidance for AI in health.

Implementation lifecycle

From a defined problem to monitored practice

  1. Define

    Specify the clinical problem, intended user, excluded use, acceptable failure thresholds and escalation pathway.

  2. Validate locally

    Test representative data and realistic edge cases before the system influences care.

  3. Pilot with oversight

    Observe how the tool changes decisions, workload and attention—not only whether its isolated answers appear correct.

  4. Monitor and revise

    Track errors, overrides, drift, subgroup effects and product changes with a route to pause use when risk changes.

Practical checklist

What an evidence-first workflow should make visible

  • The clinical question and relevant patient context.
  • Important missing data and time-critical concerns.
  • Competing hypotheses rather than a single premature conclusion.
  • The source and date of supporting evidence.
  • The limits of evidence and model knowledge.
  • What would meaningfully change the assessment.
  • Where human review and escalation are required.
  • How feedback, incidents and model changes are recorded.
Important limitation: AI outputs can be incomplete, outdated, biased or incorrect. They should not be treated as a diagnosis, directive or substitute for independent clinical assessment. Diagnify is for qualified clinicians and remains under evaluation.

Continue exploring

From principle to clinical workflow

Built for clinician review

Explore evidence-first clinical decision support

Diagnify is a decision-support environment under evaluation for qualified clinicians. It does not replace clinical judgement.