Screening results, read properly

A positive result does not mean what most people think it means.

The same test, with the same accuracy, means something completely different for two different people. What changes is not the test. It is how likely the condition was before the test was ever run. This tool works that out for type 2 diabetes screening and shows the answer as one hundred people.

Step 1

What is true of you today

Ten things you already know. No laboratory value among them. These feed a risk model fitted on 253,680 responses to the CDC's 2015 behavioural risk survey.

Weight in kg divided by height in metres, squared.

Has a health professional ever told you that you have…
And about daily life

Step 2

The test, and how good it actually is

No test is perfect, and published estimates disagree. The defaults below are editable on purpose: a tool that hides this behind a single number is hiding the part that matters.

49%

Of people who do have diabetes, the share the test catches.

96%

Of people who do not, the share correctly cleared.

Step 3

The result you are holding

Before the test

your risk from the ten answers above

After the test

How much to trust the first number

The whole chain rests on the pre-test probability. If that is wrong by a factor of two, the answer is wrong by roughly the same factor. So the question worth asking of the model is not whether it ranks people well. It is whether, when it says twenty percent, twenty percent of those people turn out to have diabetes.

Measured on 76,104 held-out responses the model never saw during fitting.
Calibration error0.013mean gap between predicted and observed, over ten equal-sized bins
Brier skill score0.173improvement over always guessing the base rate
ROC AUC0.819ranking ability, reported second because it says nothing about calibration
Base rate13.9%diabetes prevalence in the survey sample
predicted by the model observed in the data

Each row is one tenth of the held-out sample, ordered by predicted risk. The pale bar is what the model predicted; the solid bar underneath is what actually happened in that tenth. They track closely, which is the property this tool depends on and the reason it is reported before anything else.

What this is not