Medical models · clinician decision support

Medical AI.
Built for clinical reasoning.

Our clinically fine-tuned models combine medical reasoning with source-linked evidence. Three connected products support clinicians, from patient history to signed note.

Research model Diagnify-14B v0.2-CoT Fine-tuned for medicine Open-weight base · clinical adaptation Evidence at every step FINE-TUNE → RETRIEVE → GROUND → CHECK Designed for clinician review
Clinician-reviewed Evidence-grounded In testing
dx.diagnify.ai Preview
Diagnify clinical decision support interface showing a structured work-up, red flags and a ranked differential diagnosis.

Clinical reasoning, in one workspace.

02
Our models

Medical models. Clinical tasks.

We fine-tune medical models and ground their reasoning in a clinician-curated evidence base. The programme spans language, ECGs, skin images and examination findings.

Research adapter

Diagnify-14B

Medical reasoning connected to our curated evidence layer.

v0.2-CoT · access-gated · in testing
Research preview

Diagnify ECG

Reconstructs 12-lead signals from images for structured interpretation.

Clinician-reviewed second read
Training in progress

Dermatology Vision

Skin-image recognition for structured examination findings.

In training · not yet released
Training in progress

Examination Vision

Normal-versus-abnormal visual findings for clinician confirmation.

In training · not yet released
Outputs can be wrong. Clinician review is required.
03
Results

A 14B model. 99.9% in our internal evaluation.

Tested on 2,500+ complex cases, graded blind by five senior physicians. The best comparator reached 83%.

DIAGNIFY expert mode
100%
GPT-5.5
83%
Claude Opus 4.8
83%
Gemini 3.1 Pro
75%
Grok 4.20
58%
Older AI GPT-4o, 2025 study
49%
DeepSeek V4
42%
Internal retrospective evaluation; not peer-reviewed. Chart percentages are rounded. GPT-4o is a separate 2025 reference. Full comparison and limits are below.
+17 pts
above the best comparator
GPT-5.5 and Claude Opus 4.8: 83%.
~2,500
clinician-confirmed cases
De-identified cases; blinded physician grading.
0%
safety-critical misses in this cohort
Comparator range: 8–19%.
04
Head-to-head results

Tested against five AI models.

Same cases. Blinded grading. Confidence intervals shown.

99.8–100% (95% CI), internal evaluation +17 pts over the next-highest model p<0.001 vs all five comparators No safety-critical miss observed in this cohort Graded by five senior physicians
Diagnostic accuracy · 2,500+ cases Correct 95% CI vs DIAGNIFY
DIAGNIFY Medical LLM structured expert mode99.9%#199.8–100%—
GPT-5.5 OpenAI83%81.5–84.5%p<0.001
Claude Opus 4.8 Anthropic83%81.5–84.5%p<0.001
Gemini 3.1 Pro Google DeepMind75%73.3–76.7%p<0.001
Grok 4.20 xAI58%56.1–59.9%p<0.001
DeepSeek V4 DeepSeek42%40.1–43.9%p<0.001
GPT-4o historical reference · 2025 study, not re-tested49%——
Unpublished internal, retrospective evaluation run by DIAGNIFY; not peer-reviewed or externally replicated. Results do not establish clinical effectiveness or patient outcomes. Five comparators used public APIs with default settings. GPT-4o is a separate 2025 reference, not re-tested.
Study methods and subgroup results

Where the accuracy gap was biggest — rare and multi-system cases.

Pre-specified subgroups: ultra-rare, multi-system, safety-critical misses, and investigation efficiency. The margin over comparators was larger in these subgroups than in the primary outcome.

Ultra-rare diseases
450 cases · prevalence <1 in 50,000
100%
vs 28–71% across the field
Multi-system disease
880 cases · ≥3 organ systems involved
100%
vs 78–79% best comparator
Safety-critical misses
life-threatening dx left off the differential
0%
vs 8–19% across the field
Unnecessary tests
investigations ordered without discriminating value
3%
vs 12–31% across the field

How the study was run.

Blinded grading, paired testing on an identical cohort, and cases post-dating the models.

5
senior physicians, blinded
Graders across internal medicine, clinical genetics, neurology, oncology and vascular medicine; inter-rater agreement Cohen's κ = 0.94 (0.92–0.96).
p<0.001
vs each comparator
McNemar's test for paired proportions with Bonferroni correction for multiple comparisons; significant against all five models.
2024–26
post-dates the models
Cases were published after the models were built and confirmed by histopathology, genetics, imaging or specialist consensus; de-identified in accordance with the Privacy Act 1988 (Cth) and the Australian Privacy Principles.
05
Specialty coverage

Built specialty by specialty.

Green: in testing. Orange: on the roadmap.

In testing Coming soon
General Practice Emergency Medicine Behavioural Paediatrics General & Internal Medicine Cardiology Respiratory & Sleep Medicine Gastroenterology & Hepatology Endocrinology & Diabetes Nephrology Neurology Rheumatology Haematology Medical Oncology Infectious Diseases Clinical Immunology & Allergy Geriatric Medicine Palliative Medicine General Paediatrics Neonatology Psychiatry Dermatology Intensive Care Medicine Anaesthesia Pain Medicine Rehabilitation Medicine Sport & Exercise Medicine Addiction Medicine Sexual Health Medicine Occupational & Environmental Medicine Public Health Medicine Clinical Genetics Nuclear Medicine Diagnostic Radiology Clinical Pharmacology
06
Decision support

Three ways to work a case.

Type, talk or use the consultation transcript. The same reasoning engine powers each.

Interactive

Work through a case as new findings update the differential and plan.

Voice

Discuss the case, challenge the differential and review management hands-free.

Ambient scribe

Turn the consultation transcript into structured findings and decision support.

07
How it works

Reasoning you can follow.

Red flags, differential, investigations and evidence checks — on one reviewable path.

STEP 01 · SCREEN

Screen

Raise red flags and must-not-miss conditions first.

STEP 02 · REASON

Reason

Update the differential as findings arrive.

STEP 03 · WORK UP

Work up

Choose investigations that distinguish the leading possibilities.

STEP 04 · VERIFY

Check

Check sources, doses and contraindications before clinician review.

08
Application

Everyday and chronic care.

In testing for non-urgent consultations and chronic-disease reviews, including rural and remote care.

History

Structured history-taking

Capture red flags and the questions that distinguish likely diagnoses.

Self-examination

Guided self-examination

Guide defined home checks and record findings for clinician review.

Visual examination

Single-frame visual examination

Use a single camera frame at the examination step.

Reasoning

Same reasoning architecture

Review the differential, work-up and assessment together.

Scope

In testing. Clinician oversight required.

09
Before the appointment

History ready before the appointment.

Patients answer an adaptive history in their language. Clinicians receive a structured summary.

→

History only

No diagnosis, triage category or treatment advice.

→

Clinician confirms

Review and correct the history during the consultation.

→

Emergency guidance

Anyone acutely unwell is directed to call 000.

10
Privacy & data residency

Your patients' data stays in Australia.

✓

Australian GPU compute

GPU inference runs in Sydney.

✓

Sydney data residency

Patient data is stored in Sydney and stays in Australia.

✓

Encrypted records

Encrypted in transit and at rest.

Important — please read

Decision support, in testing. A registered practitioner must verify every output. Benchmark results are not clinical validation or regulatory approval. In an emergency, call 000.