Diagnify reasoning workflow

Clinical reasoning AI

Diagnify combines a clinically fine-tuned open-weight model with a separate evidence and guideline retrieval layer. It structures a case, surfaces alternatives and returns evidence for a qualified clinician to review.

50M+ consultation records 50 years represented Human feedback during fine-tuning

Two-layer foundation

Clinical adaptation first. Current evidence at the point of use.

Training teaches the model clinical patterns; retrieval connects each live case to a separately maintained evidence corpus.

  1. Curated clinical histories

    Diagnify currently describes its internal corpus as more than 50 million historical consultation records spanning 50 years. The counting method, provenance and coverage limits are tracked in the research register.

  2. Selected open-weight model

    A suitable base model is clinically fine-tuned rather than presented as a general-purpose model with a medical prompt.

  3. Training-time human feedback

    Expert reviewers identify weak reasoning, correct responses and feed those corrections back into model development.

  4. Separate evidence retrieval

    At use time, the model retrieves from a curated, versioned guideline and evidence library so consequential claims can be inspected.

Distinct controls: human feedback during fine-tuning improves model behaviour; clinician oversight during use remains necessary because neither fine-tuning nor retrieval eliminates error.

Core workflow

Screen, reason, investigate, verify

A structured cycle reduces the chance that an AI-generated suggestion is mistaken for a conclusion. The clinician controls every transition.

  1. Screen

    Identify instability, time-critical conditions, safeguarding concerns and information that cannot safely wait for a complete work-up.

  2. Reason

    Build a concise problem representation, generate competing hypotheses and identify findings that support, weaken or fail to explain each one.

  3. Investigate

    Select history, examination or tests only when they can meaningfully change probability, management or the need to escalate.

  4. Verify

    Check consequential statements against appropriate sources, reconcile contradictions and document the clinician’s independent assessment.

The cycle repeats: new findings should update the problem representation and differential. A model output is a prompt for review, not an endpoint.

Reasoning structure

What useful AI support should reveal

Problem representation

A focused summary of relevant features, chronology, context and severity—without flattening important ambiguity.

Competing hypotheses

Several plausible explanations, including dangerous alternatives and non-disease explanations where relevant.

Discriminating findings

The observations that would most change relative probabilities, rather than a long undifferentiated list of questions and tests.

Missing information

Gaps that limit confidence, with no inference that absent documentation means a finding is absent.

Uncertainty

What remains unknown, how sensitive the assessment is to assumptions, and when uncertainty itself warrants review.

Evidence traceability

Links from consequential claims to inspectable evidence, including source date, population and limitations.

Bayesian updating

Revise probability as information arrives

Bayesian reasoning begins with a pre-test probability, expresses it as odds, applies the likelihood ratio of a finding, and converts the result back to a post-test probability.

The calculation

Pre-test odds = probability ÷ (1 − probability)

Post-test odds = pre-test odds × likelihood ratio

Post-test probability = post-test odds ÷ (1 + post-test odds)

Calculate post-test probability with Diagnify’s educational Bayesian tool.

The clinical caveats

  • Pre-test probability must fit the patient and setting.
  • A likelihood ratio may not transport across populations or methods.
  • Correlated findings should not be treated as independent without justification.
  • A probability does not set a universal treatment or testing threshold.

Two human-control layers

Review during training. Oversight during every clinical use.

Training-time human feedback

  • Reviewers assess clinical relevance and reasoning quality.
  • Incorrect, incomplete and unsafe responses are corrected.
  • Disagreement and difficult cases inform further evaluation.
  • Human feedback improves behaviour but does not guarantee accuracy.

Clinician oversight at use time

  • Form an independent problem representation and address urgency first.
  • Check for anchoring on the model’s ordering or wording.
  • Inspect retrieved sources and recalculate consequential estimates.
  • Retain responsibility for assessment, escalation and management.
Not autonomous care: clinical reasoning AI can hallucinate, omit crucial diagnoses, misread context and present uncertainty poorly. Diagnify is under evaluation for qualified clinicians and is not intended for emergencies or patient self-diagnosis.

Related guidance

Explore the reasoning stack

Two ways to work with Diagnify

Use the clinical product or build on the same foundation

Qualified clinicians can use Diagnify’s reviewed workflow. Medical AI teams can request a private, dedicated model and evidence API deployment for their organisation.