Diagnify-14B
Medical reasoning connected to our curated evidence layer.
Our clinically fine-tuned models combine medical reasoning with source-linked evidence. Three connected products support clinicians, from patient history to signed note.
Research model Diagnify-14B v0.2-CoT Fine-tuned for medicine Open-weight base · clinical adaptation Evidence at every step FINE-TUNE → RETRIEVE → GROUND → CHECK Designed for clinician reviewClinical reasoning, in one workspace.
Take the history, work up the case, and sign the record — together or as separate tools.
Evidence-linked differentials, investigations and plans that update as findings arrive.
Voice history before the appointment, in the patient’s language.
Intake, scribe and decision support inside your practice record.
We fine-tune medical models and ground their reasoning in a clinician-curated evidence base. The programme spans language, ECGs, skin images and examination findings.
Medical reasoning connected to our curated evidence layer.
Reconstructs 12-lead signals from images for structured interpretation.
Skin-image recognition for structured examination findings.
Normal-versus-abnormal visual findings for clinician confirmation.
Tested on 2,500+ complex cases, graded blind by five senior physicians. The best comparator reached 83%.
Same cases. Blinded grading. Confidence intervals shown.
| Diagnostic accuracy · 2,500+ cases | Correct | 95% CI | vs DIAGNIFY |
|---|---|---|---|
| DIAGNIFY Medical LLM structured expert mode | 99.9%#1 | 99.8–100% | — |
| GPT-5.5 OpenAI | 83% | 81.5–84.5% | p<0.001 |
| Claude Opus 4.8 Anthropic | 83% | 81.5–84.5% | p<0.001 |
| Gemini 3.1 Pro Google DeepMind | 75% | 73.3–76.7% | p<0.001 |
| Grok 4.20 xAI | 58% | 56.1–59.9% | p<0.001 |
| DeepSeek V4 DeepSeek | 42% | 40.1–43.9% | p<0.001 |
| GPT-4o historical reference · 2025 study, not re-tested | 49% | — | — |
Pre-specified subgroups: ultra-rare, multi-system, safety-critical misses, and investigation efficiency. The margin over comparators was larger in these subgroups than in the primary outcome.
Blinded grading, paired testing on an identical cohort, and cases post-dating the models.
Green: in testing. Orange: on the roadmap.
Type, talk or use the consultation transcript. The same reasoning engine powers each.
Work through a case as new findings update the differential and plan.
Discuss the case, challenge the differential and review management hands-free.
Turn the consultation transcript into structured findings and decision support.
Red flags, differential, investigations and evidence checks — on one reviewable path.
Raise red flags and must-not-miss conditions first.
Update the differential as findings arrive.
Choose investigations that distinguish the leading possibilities.
Check sources, doses and contraindications before clinician review.
In testing for non-urgent consultations and chronic-disease reviews, including rural and remote care.
Capture red flags and the questions that distinguish likely diagnoses.
Guide defined home checks and record findings for clinician review.
Use a single camera frame at the examination step.
Review the differential, work-up and assessment together.
In testing. Clinician oversight required.
Patients answer an adaptive history in their language. Clinicians receive a structured summary.
No diagnosis, triage category or treatment advice.
Review and correct the history during the consultation.
Anyone acutely unwell is directed to call 000.
GPU inference runs in Sydney.
Patient data is stored in Sydney and stays in Australia.
Encrypted in transit and at rest.
Decision support, in testing. A registered practitioner must verify every output. Benchmark results are not clinical validation or regulatory approval. In an emergency, call 000.