Model development and transparency

How Diagnify fine-tunes and evaluates medical models.

Diagnify’s current language model was not trained from scratch. Diagnify selected a licensed open-weight base model, applied medical fine-tuning, and evaluates the resulting model separately from the inference-time evidence-retrieval system used in Diagnify.

Selected open-weight base Medical fine-tuning Access-gated research adapter

Current documentation status: the exact base-model identifier and checksum, complete licence lineage, training-data statement, human-feedback report and benchmark report are not yet published on this page. The process and performance claims therefore cannot yet be independently assessed. Current evaluation work is internal and retrospective, not prospective clinical validation or evidence of improved patient outcomes.

Development record

A selected base model, a clinical fine-tune, then a separate evidence layer.

These stages should not be collapsed into the claim that Diagnify trained a foundation model from scratch.

Select and license the base

Choose an open-weight base model for defined technical and clinical requirements. Record its exact name, version, checksum, licence and restrictions.

Prepare the clinical corpus

Document source rights, counting rules, de-identification, inclusion and exclusion, quality controls, representation gaps and safeguards against evaluation leakage.

Fine-tune with human feedback

Apply medical domain adaptation and a documented reviewer-feedback process. Reviewer qualifications, rubric, sampling, double review and adjudication need their own report.

Evaluate the frozen model

Test a named model version on held-out cases under a prespecified protocol, then report uncertainty, subgroup performance, errors and protocol deviations.

Evaluate retrieval separately

At inference, guideline retrieval remains outside the learned weights. Retrieval recall, source currency, quotation fidelity and citation support require separate evaluation.

Company-supplied corpus description

What “50M+ consultations across 50+ years” does—and does not—say.

Diagnify describes its internal clinical corpus as more than 50 million consultation records spanning more than 50 years and sourced from multiple regions. These figures are supplied by Diagnify and have not been independently verified on this page. “Consultation” does not necessarily mean a unique patient, and multi-region sourcing does not establish global representativeness.

Count and coverage still need definition

  • Snapshot date and corpus version
  • Counting unit and duplicate handling
  • Start and end years, not duration alone
  • Patients, encounters, notes and examples distinguished

Provenance and rights need a record

  • Source categories, regions and proportions
  • Languages, settings and specialties represented
  • Permission, licence, legal or ethics basis
  • De-identification method and residual risks

Quality needs measurable controls

  • Eligibility, exclusions and sampling
  • Missingness, label quality and temporal drift
  • Subgroup gaps and known biases
  • Train-test contamination and leakage checks

Publication gate: the 50M+ and 50+ year description should be interpreted as an internal corpus statement until a dated data statement supplies these definitions, permissions, distributions and limitations. Volume alone is not evidence of clinical quality.

Documentation register

What is public, internal or still planned.

Status is attached to each artefact so an internal description cannot be mistaken for independently inspectable documentation.

Base model and adapter record

The current Hugging Face asset is an access-gated research adapter. A complete record should identify the base model, checksum, licence lineage, fine-tuning method, intended use and known limitations.

Status: access-gated adapter; complete base-model lineage is not yet documented here.

Training-data statement

Corpus version, unit of count, provenance, rights, temporal and regional distribution, preparation, de-identification, exclusions, representation limits and leakage controls.

Status: company-supplied high-level description only; source-level documentation is not public here.

Human-feedback report

Reviewer number and credentials, training, rubric version, sampled examples, double-review rate, agreement, adjudication, quality checks, dates, funding and conflicts.

Status: required for assessment; not yet published on this page.

Evaluation protocol and report

Prespecified question, cohort, endpoints, exact comparators and settings, grading, analysis, uncertainty, subgroups, errors, missing data, deviations and reproducibility materials.

Status: current work is described as internal and retrospective; no complete report is linked here.

Retrieval evaluation

Source coverage and currency, retrieval recall, jurisdiction matching, quotation fidelity, citation support, conflicting guidance, failure analysis and update performance.

Status: retrieval is described separately from model evaluation; no public report is linked here.

Change log and corrections

Material model, data, protocol, retrieval and report changes should be dated, explained and linked to prior versions, negative findings and known errors.

Status: format defined; populated version history remains planned.

No universal superiority claim: until a versioned comparative report is public, this page does not claim that Diagnify generally surpasses larger or open-weight models. Any future comparison must name the exact systems, test conditions, endpoint, uncertainty and limitations, and should be read as a result for that evaluation—not as proof of clinical benefit.

Interpretation guide

Different evidence answers different questions.

These labels prevent one kind of evaluation from implying a stronger claim than its design supports.

Interpretation of research and validation labels
Label What it can examine What it does not establish by itself
Internal retrospective evaluation Performance on previously assembled cases under a documented test procedure Independent reproducibility, prospective performance or clinical benefit
Independent replication Whether results can be reproduced by investigators outside the development team Safe and effective performance in routine clinical workflow
Prospective validation Performance on cases or data collected after a protocol is set Improved patient outcomes unless outcomes and study design address that claim
Clinical outcomes study Effects on prespecified care, safety or patient outcomes in a defined setting Generalisation to populations, settings or uses outside the study
Minimum disclosure

A benchmark needs more than a headline result.

Before an internal evaluation is suitable for public interpretation, its method and boundaries need to be visible.

System identity

  • Exact model and application version
  • Prompting, retrieval and tool configuration
  • Evaluation date and frozen settings

Evaluation design

  • Prespecified question and endpoints
  • Case provenance and eligibility rules
  • Comparators, grading and analysis method

Interpretive limits

  • Uncertainty and missing data
  • Subgroups, bias and error analysis
  • Generalisability and unresolved risks
Release discipline

From internal work to a citable record.

Publication status should progress only when the corresponding material is actually available.

Define

State the research question, intended claim, protocol, system version and analysis before interpreting results.

Run and audit

Preserve outputs, errors, deviations and exclusions. Separate exploratory analysis from prespecified analysis.

Report

Publish methods, results, uncertainty, limitations and a version identifier together. Clearly label internal authorship and retrospective design.

Invite verification

Provide reproducibility material where permissions allow, distinguish independent work from internal work, and link corrections without erasing the record.

Change log format

Every material change gets context.

Research pages will use a consistent record rather than silently replacing previous claims or methods.

Each entry records

  • Date and version affected
  • Resource, model or protocol changed
  • Plain-language summary and reason
  • Whether results or interpretation changed

Each status distinguishes

  • Draft or internal material
  • Publicly released documentation
  • Corrected or superseded material
  • Independent work linked from this hub