What Medical AI Doesn't Know

Missing-data benchmark

What happens to model performance when part of a clinical record is missing?

I compare three model families on two public datasets under controlled missingness. In the explorer, I report changes in ranking, probability error, and calibration, together with the amount of the full feature matrix that was actually removed.

Illustrative partial record6 fields
Age
61
Resting blood pressure
128 / 84
Cholesterol panel
Missing
Resting ECG
Normal
Max heart rate
Missing
Chest pain type
Atypical
2 of 6 fields unavailable
I use this record only as an illustration. My experiments use benchmark feature matrices, not patient records entered through the site.

Why the denominator matters

A 50% target-feature rate does not remove half of the record.

In my primary comparison, I allow missingness in five of 30 WDBC features or three of 13 Statlog Heart features. The requested rate applies inside those target sets.

WDBC

8.3%

Approximate whole-matrix missingness at the requested 50% target-feature rate.

Statlog Heart

11.5%

Approximate whole-matrix missingness at the same requested target-feature rate.

Principal result

Scope-matched mechanisms produced small ROC-AUC changes. The all-feature stress test was much harsher.

I therefore analyze MCAR, MAR, and MNAR on shared target sets, then report MCAR across the full feature matrix as a separate experiment.

Study at a glance

Datasets

2

WDBC and Statlog Heart provide different sample sizes, feature structures, and clean reference performance.

Models

3

I evaluate logistic regression, random forest, and histogram-based gradient boosting with shared masks and outer splits.

Mechanisms

3

I implement MCAR, MAR, and MNAR as controlled simulation rules on matched target features.

Requested rates

4

I pair each requested target-feature rate with achieved target-set and whole-matrix rates.

Largest scope-matched ROC-AUC decline

-0.006

Logistic Regression on Statlog Heart classification, using MCAR at a requested 50% target-feature rate. This is a conditional rate, not a whole-matrix percentage.

How I produced the results

  1. 01

    I load two public clinical benchmark datasets

  2. 02

    I apply matched MCAR, MAR, and MNAR target-feature masks

  3. 03

    I evaluate three model families on the same outer splits

  4. 04

    I compare discrimination, probability error, and calibration

  5. 05

    I export the saved results to a static web artifact

Disclaimer

Research and education only

I use this site to report controlled experiments on public benchmark data. It is not medical advice, a diagnostic tool, or a basis for personal health decisions.

  • No symptom assessment or chatbot
  • No medical-data upload or storage
  • No diagnosis or treatment guidance