Insights

Evaluating a pilot honestly

Most pilots in this field cannot fail. That is what is wrong with them.

A pilot that cannot produce a negative result is not a pilot. It is a demonstration, and there is a shelf of them in every health department in the country.

The four things that make a pilot uninformative

  1. A convenience population. Enrolling whoever is easiest to reach guarantees the program works, because the people it would fail were never in it. A pilot population should be defined by a rule that includes the hard cases.
  2. No comparison. This one is acute right now: opioid-involved overdose deaths fell nationally from 82,851 in 2022 to 55,007 in 2024. Any uncontrolled program running in that window will report a large improvement it had nothing to do with.
  3. Too short. Overdose is a low-frequency event. A three-month pilot in a small population cannot detect a change in it, and reporting one is reporting noise.
  4. Measures chosen after the data exists. Which turns any dataset into a success.

What to fix them with

  • A stepped-wedge design, where sites start at staggered times. Everyone gets the program eventually, which is usually what makes it acceptable, and the staggering supplies the comparison.
  • Pre-registered measures, including the safety ones from measurement and evaluation.
  • A stopping rule agreed in advance — what result would halt this early. A program with no stopping rule cannot fail, and therefore cannot succeed either.
  • An independent evaluator with access to the underlying data, not just to the vendor’s summary.
  • Publication regardless of result.

The measures that should worry you

A pilot reporting only favorable measures is not reporting. The ones to insist on are the ones that would embarrass the program: patients leaving care, therapy gaps, enrollment skewed toward the easy cases, and differential outcomes by race, insurance status or geography. That last one will not announce itself — it has to be looked for on purpose.

Why we want this

Partly integrity, and partly self-interest. A program that has only ever been evaluated by itself is worth nothing in a legislative hearing or a competitive procurement. The first agency to evaluate this properly, and publish, does the program a larger favor than a decade of favorable case studies.

More: pilot design.

Sources

Every figure on this page is traceable to the source listed here.

  • National Center for Health Statistics. VSRR Provisional Drug Overdose Death Counts (dataset xkb8-kh2a), 12-month-ending counts for the United States. Data accessed 20 August 2026. View source.