Insights

Base rates and outlier flags

When the behavior you are screening for is rare, most of what you flag will be something else. This is arithmetic, not a criticism of anyone’s analytics.

Every outlier-detection system in prescribing analytics runs into the same problem, and it is not a problem of technique. It is a property of screening for rare things.

The rate to start from

The best available estimate of concentrated, potentially diversionary prescription acquisition analyzed 146.1 million opioid prescriptions dispensed in 2008 across pharmacies representing 76% of US retail volume. The extreme outlier group — averaging 32 prescriptions from 10 prescribers — was 0.7% of purchasers, accounting for 1.9% of prescriptions and about 4% of the weighed amount.

The authors were explicit that even this group cannot be assumed to be diverting: “Very few of these patients can be classified with certainty as diverting drugs for nonmedical purposes.”

What that does to a flag

Take a generous screening instrument: 90% sensitivity and 95% specificity, applied to a population where 0.7% are true positives. In 100,000 people that is 700 true cases. The instrument catches 630 of them — and also flags 5% of the 99,300 who are not, which is 4,965 people.

So of 5,595 flags, 630 are correct. About 11%. Nine out of ten flagged people are not the thing you were looking for, with a test that is 95% specific.

Push specificity to 99% and you still get 993 false positives against 630 true ones — 39% precision, with a test that is wrong about one time in a hundred.

Three consequences

  1. Tightening the threshold does not fix it. It trades false positives for false negatives along the same curve. The arithmetic is driven by the base rate, not the cutoff.
  2. The absolute harm scales with the population. Which is why the size of the chronic pain population matters — about 50 million US adults, with 19.6 million experiencing high-impact chronic pain. A small false-positive rate against a denominator like that is a large number of people.
  3. The only real fix is a second, independent source of information. Something that resolves the ambiguity rather than sharpening the same signal. Diagnosis and treatment intent are exactly that, which is the entire argument of counting without context.

This is not an argument against screening

Screening for rare conditions is worthwhile in medicine all the time — with a confirmatory second step, and with everyone understanding that a positive screen is a question rather than an answer. The failure mode is treating a screen as a finding. In prescribing analytics, that is what happens when a flag becomes an action without a human who can see the clinical picture.

More: ProviderSynch · rule discovery.

Sources

Every figure on this page is traceable to the source listed here.

  • McDonald DC, Carlson KE. Estimating the prevalence of opioid diversion by “doctor shoppers” in the United States. PLoS One. 2013;8(7):e69241. PMID 23874923. View source.
  • Dahlhamer J, Lucas J, Zelaya C, Nahin R, Mackey S, DeBar L, Kerns R, Von Korff M, Porter L, Helmick C. Prevalence of chronic pain and high-impact chronic pain among adults — United States, 2016. MMWR Morb Mortal Wkly Rep. 2018;67(36):1001–1006. PMID 30212442. View source.