I have watched many hospitals make a concerted effort to remove central lines when patients no longer needed them. It was a deliberate and intentional effort, and it worked. Line days came down. Fewer patients had a CVC and the risk that comes with central access decreased.
Then the central line infection rate went up.
The care had not gotten worse. It had gotten better. That rate, CLABSI in the reports, is infections per thousand line days. The patients who still had lines were the ones who could not do without them. They were sicker, they stayed longer, and so did their lines. Their infections stayed in the numerator. The lower-risk line days, the ones that rarely turned into an infection, left the denominator. What remained was divided by far fewer days.
This is common. It is common enough that infection prevention researchers have made the case for a metric that accounts for how many patients have lines at all, not only for what happens once they do, because reducing unnecessary device use may select a higher-risk population that current risk adjustment does not account for. Their phrase for what follows is “a paradoxical increase” in the standardized infection ratio “for facilities that may be high performers.” That ratio is the one a hospital is publicly reported and paid on.
Those numbers moved without anything about the care getting worse. Every quality number a clinical leader will be handed can do the same thing, in one direction or the other.
I have a stake in this argument. I help lead a company that is judged on numbers like these, and when one of ours goes the wrong way, “the counting changed” is the first explanation I reach for. Some of ours are in play right now. I will not say which, or which way, so hold me to the test at the end of this essay. It is the one I have to pass.
Why a CLABSI rate goes up when the care improves
There are three ways the counting moves a number without the clinical care changing. The definition of what counts changes. How hard we look, and who does the judging, changes. Or the population left in the denominator changes. The central line story is the third. It is the easiest to miss, because the change that caused it was the improvement itself.
The first two are better documented, and the best-documented example runs in the flattering direction.
When the definition changes
In January 2015 the national surveillance definition for catheter-associated urinary tract infection changed. Yeast stopped counting. So did cultures below a hundred thousand colony-forming units per milliliter. By the middle of that year the national ratio of observed to predicted infections had fallen to 0.55, an estimated forty-five percent reduction in a single year.
CDC did not call that progress. It told hospitals to adjust their older numbers before claiming a prevention effect, and warned that hospitals “may appear to have met the HHS goal even if they still require further CAUTI prevention efforts.”
In 137 intensive care units measured the year before the change and the year after, the urinary infection rate fell 44.8 percent. The central line infection rate in the same units rose 27 percent. A fungal bloodstream infection that would once have been attributed to the urine culture now had no urinary tract infection to attach to, and in a patient with a central line it was counted against the line. Central line infections from Candida rose 91 percent.
Same units, one year apart. A number that falls gets a slide in the board deck. A number that rises gets a meeting. Both should invite the question of how it was counted, but rarely do.
It is not only CAUTI. Applied to the same 5,804 surgical wounds, one accepted definition found 19.2 percent of them infected and another found 6.8 percent. A definition alone can move a rate almost threefold. Same wounds, same patients, same hospitals.
How hard you look
How hard you look has an equal impact. Across English hospitals, knee replacement infection rates were 4.1 percent where follow-up after discharge was thorough and 1.5 percent where it was not. The authors wrote that hospitals doing the better surveillance will be penalized.
A Swiss program covering 164 hospitals and 187,501 operations found that the share of infections detected only after the patient went home ran from about a fifth after colon surgery to more than nine in ten after knee replacement. A program that stopped looking at discharge would have missed more than nine in ten of the knee infections. The authors titled the paper for the finding: Who Seeks Shall Find. Their conclusion was that intensive surveillance after discharge creates artificial differences between programs. As it turns out, the harder you look, the more you find.
I have seen the same thing in my own specialty, across a number of implementations in different systems over a number of years. When an anesthesia group moves from paper records to an anesthesia information management system, its documented rates of hypotension, hypothermia and difficult airway go up, in every one I have seen. The system records blood pressure and temperature from the monitor on a fixed interval for the whole case, and it asks for the number of intubation attempts and the view grade in fields of their own. A clinician charting by hand in the middle of a case records what there was time to write down, and adverse events are counted off what the clinician reports rather than what is measured against a standard definition. I have compared sites across that line and taken longer than I should have to see how much of the gap was the record. Some of it may be the patients. Until both sites are measured with the same kind of record, there is no way to know how much.
Who calls it an infection
Deciding which cases count sounds clerical. It is not.
When surgeons and infection preventionists in ten countries were given the same twelve written cases, agreement among the surgeons was 0.24 on a scale where 1.0 is perfect and zero is chance. The infection preventionists reached 0.41. After the surgeons read the formal definitions, their agreement fell to 0.09.
Consistency is not accuracy either. In seven Dutch hospitals, the people doing surveillance on colorectal surgery agreed with each other about 95 percent of the time. They still disagreed with an outside panel on roughly one case in eight.
Those judgments become a person’s number. It lands in an evaluation file with a name on it, or in a case review with the clinician who last cared for the patient, and that clinician rarely hears who decided which of their cases counted, or whether that changed.
The test
Before you ask what happened to the patients, ask what happened to the counting.
That is a question with an answer, and the answer is three dates. The last time the definition changed. The last time the way we capture it changed: a new record system, a new reporting form, a new surveillance protocol, a new person or committee deciding which cases count. The last time the population it counts changed: a line-removal campaign, a new discharge pattern, a service line that opened or closed.
Put those three dates on the slide next to any trend line. Better yet, mark these changes on the trend line directly. If one of them falls inside the window the line covers, the line cannot be read as care until someone has separated the counting from the care. If none of them does, the counting stops being an explanation. Infection prevention and quality analytics can usually do that separation if they are asked. Ask for two things: the older period recounted under the current rules, and the denominator plotted beside the rate. They are rarely asked, because a number that improved does not prompt the question and a number that worsened prompts a different one.
When one of our own numbers moves, in either direction, I owe the three dates, and if one falls inside the window, the older period recounted under the current rules, before I credit the care or blame the counting. Without that, the counting is not my explanation, and I do not offer it.
The people who pulled those central lines did the right thing. The number said they had made things worse. Both were true on paper, and only one was true for the patients.



The anesthesia example is so true, David! The minute a group goes from paper to an AIMS, the documented rates of hypotension and difficult airway jump, and it's the monitor doing the recording rather than the patients getting sicker. And the three dates is such a practical ask — I'm stealing it.