ROC Curve & AUC

Draws the ROC curve, reports the area under it with a DeLong interval, and lists every cutoff with its sensitivity and specificity.Paste your measurements and reference standard (1 or 0) below.

This tool describes how a measurement performed across a set of samples. Do not use it to decide anything about an individual patient.

One sample per line: the measurement first, then 1 for the condition present and 0 for absent, separated by a tab, comma or spaces. A header row is skipped.

How to use it

  1. Paste two columns: the measurement, and the reference standard as 1 for the condition present and 0 for absent. A header row is skipped.
  2. The curve, the area and the interval appear as you type. The direction is read off the data — a marker that falls in the condition is handled without a setting, and the result says which way it was read.
  3. The Youden cutoff is the criterion where sensitivity plus specificity is largest. It is marked on the curve and highlighted in the table.
  4. Scroll the table for the whole ladder of criteria. Picking one by a requirement — the lowest criterion with at least 95% sensitivity, say — is what the table is for.
  5. Take the criterion you chose to the 2×2 tool for predictive values, which depend on prevalence and so are not on this page.

💡 This tool describes how a measurement performed across a set of samples. Do not use it to decide anything about an individual patient.

Method

The area under a ROC curve is the Mann–Whitney statistic: the proportion of (positive, negative) pairs in which the positive scored higher, counting a tie as one half. It does not depend on prevalence, and neither do sensitivity and specificity, which is why this page reports those three and leaves the predictive values to the 2×2 tool.

The standard error is the method of DeLong, DeLong and Clarke-Pearson (1988), which MedCalc names as its recommended one. The interval is the area plus and minus 1.96 standard errors, and it is cut off at 0 and 1 when it runs past them — the page says so when it has.

The cutoff is the criterion that maximises Youden's J, sensitivity + specificity − 1, over the criteria that were actually observed. When several share the maximum the page says how many, because the choice between them is then arbitrary.

  • ⛔ There is no grading of the area on this page. Bands such as “above 0.7 is acceptable” have no primary source behind them, and printing one would be making a judgement on the reader's behalf.
  • ⛔ A cutoff maximising Youden's J weighs a false positive and a false negative the same. When they are not the same in the setting the measurement is used in, the maximum is the wrong criterion and the table is where the right one is found.
  • ⚠️ The direction is chosen by the data: if the area comes out below a half with the higher-is-positive reading, the marker runs the other way and the whole picture is that one reflected. The result states which reading it used.
  • ⚠️ A criterion chosen on the same samples that estimated the curve is optimistic. Reported performance at a cutoff picked this way needs a separate set to stand up.
  • ⚠️ Ties count as half a pair. A measurement reported to one decimal makes far more ties than one reported to three, and that alone moves the area.

FAQs

Often used together