A surgeon sat across from me to talk about a surgical site infection rate that was higher than peers, and a standardized infection ratio that matched. A ratio is observed infections against what a risk model predicted. The numbers came from our infection prevention team, who had produced them carefully and correctly.
I opened by acknowledging the case mix difference. The practice ran more heavily to cancer than the group it was being compared against, and I said so. I said the cases ran longer. I said the patients arrived with more coexisting disease. I said there was more transfusion. I said all of it before the surgeon had to.
Every one of those things was true.
Then I said the number was nuanced.
I have thought about that word for a long time since. It is the most useful word in the executive vocabulary, and the reason it is useful is that it concedes a problem without naming one. It sounds like candor. It cannot be checked. It let me be honest and unhelpful at the same time, and the person across the table walked out frustrated and with nothing that could be used to improve.
I have described that meeting since as giving someone the benefit of the doubt. What it actually did was close the subject.
And the case mix was never the surgeon’s missing piece. Anyone operating on that population knows their patients are not the comparison group’s. What the surgeon did not have was the instrument — what a standardized infection ratio even is, let alone what had gone into that one. I had the vocabulary, and I had someone else’s analysis. What I had not done was open the model and read what was inside it, and I was the only person at that table who could have. Arriving with the mitigation already assembled meant the surgeon never had to make the case, and so never received the one thing that would have made the case without me in the room.
What the risk adjustment accounts for, and what it does not
Everything I said in that room was accurate. None of it was usable.
What I did not say was what the number was built from. I came to the table with a value and detailed analysis done by another team. Carrying a number and understanding it are separate activities, and I had done the first while sounding like I had done the second.
The risk adjustment behind that ratio, under the inpatient model in force before it changed, carried five things. Whether the patient had diabetes. An ASA physical status class. Whether body mass index was thirty or above. Age. And a flag for whether the hospital is enrolled as an oncology hospital.
That last one is a fact about the facility. It is not a fact about the patient on the table.
Cancer is not in the risk adjustment model. Not as a diagnosis, not as an indication, not anywhere. The effect is published and it is not small. Gynecologic cancer carries an independent adjusted odds ratio of 1.54 for surgical site infection after hysterectomy across 66,001 of those operations at 166 hospitals, and a pooled 1.49 across 681,695 patients in twenty-three studies.
The model has been rebuilt since that conversation. It now includes nine variables instead of five, including how long the operation took. It is a better and more specific model. Cancer still is not one of the model variables.
So I was right about the case mix, and the instrument could not represent it. Not because anyone decided cancer should not count. Because a hospital enrollment flag was used where a patient variable belonged, and because that flag adjusts every surgeon in an enrolled building identically, which means it cannot tell apart the two practices the number was being used to compare.
Now grade the other things I said in that room, which I did not do at the time.
Comorbidity was already in the number. Diabetes, physical status, body mass index and age are four of the five variables, so I gave credit for something the model had counted — the adjustment had already run in the surgeon’s favor before I opened my mouth. Transfusion has never been a covariate, and does not hold up well in the literature either: null across five randomized trials, and dropping out on multivariable adjustment in the hysterectomy analysis above.
The longer operating times were the exception. Operative duration had no term in the model then in force, which made it the one thing I said that the instrument could not already answer. I said it once and moved on.
So of the three: one already counted, one unsupported, one real and unpressed. None of that was clear to me at the time, which is the point.
A group working through 6,142 procedures at one public health system added clinical variables to the CMS set and improved the model’s discrimination from .675 to .735, calling the current risk assessment overly simplistic. In their own matched comparison the uterine cancer proportions did not differ. I include that second sentence because leaving it out would be the thing this essay is about.
The floor is not the standard
NHSN will not calculate a ratio at all when fewer than one infection is predicted, in order to enforce precision of the estimate and comparisons to national data.
That is a floor. It decides whether a number may be printed.
It is not the standard for judging a person. The rule of thumb statisticians use is ninety percent reliability before drawing a conclusion about an individual. Seventy to eighty is considered acceptable for a group.
No single clinician’s volume comes near it. In a 345-surgeon collaborative, eighty-six percent of the variation between surgeons was measurement noise. 168 cases across three years were required to reach even 0.7. One surgeon in 345 had the volume. That was all complications after colectomy rather than infection alone, and the arithmetic gets worse as the outcome gets rarer. At the hospital level — hospitals, not people — ninety-four cases were needed for the same threshold on superficial infection after colon resection, and about half of hospitals had them.
Most of the distance between any two surgeons in that collaborative was due to chance.
The number I walked into that room with passed the only test anyone had run on it. That test was never built to answer the question I was using it to answer.
The two numbers were the same number
Here is the part I understood last, and it is the part that should have been said first.
Infection prevention could hand me a ratio for that surgeon because it was the same ratio as the facility’s. The reported volume for that procedure, over that period, was effectively one practice’s.
The measure CMS publishes and pays on counts only the abdominal hysterectomies where the patient stays overnight. NHSN calls a procedure outpatient when admission and discharge fall on the same calendar day, and every one of those is excluded from the measure that feeds public reporting. The test is the discharge, not the room: a laparoscopic case in the hospital’s own operating room, home that afternoon, is outside the number. Two thirds of hysterectomies were already being done that way by 2013.
So the denominator was not small at random. What decides whether a hysterectomy stays in the measure is whether the patient stays overnight, and what usually decides that is how large the operation was. The cancer operations were the larger ones. In this practice more of them were open rather than laparoscopic, longer, more involved, and more likely to keep someone in a hospital bed overnight. The simpler cases went home the same afternoon and were never measured.
The reporting rule does not sample hysterectomy. It samples the hard end of hysterectomy, and then scores what it finds with a model that has no term for cancer — and, under the version in force then, no term for the approach or for how long the operation took. The model has since added both of those. It still has not added cancer.
An inclusion rule is a sampling decision, and it is almost never independent of the thing being sampled.
That is the loop. The case mix the instrument could not adjust for is the same case mix that decided who appeared in the instrument at all.
The exclusion is not a rounding error either. At one hospital, reportable abdominal hysterectomies fell from 360 in 2017 to 234 in 2021, as same-day discharge went from three percent of its cases to forty-nine. Predicted infections are the actual denominator of the ratio, and they fell thirty percent across those four years. The hospital’s infection ratio rose forty-two percent for every infection it recorded. In 2021 alone, excluding those cases cut predicted infections forty-seven percent and raised the ratio ninety percent.
Here it shrank until the reported population was effectively one practice. The facility’s number and the person’s number were the same, and a rule about what time the patient went home decided how large it was.
The reliability arithmetic above says a hospital needs ninety-four cases before its own rate means much. That test went unrun here, because on paper this was a facility measure, and facility measures are assumed to be large enough — an assumption built on research done at facility scale. It stopped being true the moment a reporting rule shrank the facility to one practice, and no step in the process is designed to notice that.
And this is the recommended practice
The surveillance manual says the ratio can be calculated for specific surgeons. No caution follows it. The 2022 infection prevention compendium goes further and recommends the practice: audit routinely, give confidential feedback of rates and ratios to individual surgeons and to chiefs, benchmark anonymously among peers. It carries one measurement caution, that the one-predicted-infection rule may be harder to satisfy for small surgical programs. That is a note about whether the number can be produced, not about whether it can tell two clinicians apart.
Everyone involved is doing what the guidance says. That is why nothing corrects it.
The asymmetry underneath is the whole thing.
The payment consequence that got the research belongs to the hospital. Six measures weighted equally, colon and abdominal hysterectomy infections combined into one of them, the worst quartile losing one percent of its Medicare fee-for-service payments. That is exactly the level at which the reliability work was done.
The career consequence belongs to a person. Ongoing professional practice evaluation, at an interval that cannot exceed twelve months, is what carries a measure to an individual file. What it usually produces is not a privilege action but a focused professional practice evaluation — the chart audit, sometimes the required-presence proctoring. Further out, an adverse action lasting longer than thirty days is reportable by statute to the national practitioner data bank, and it stays on the record.
That last one is the ceiling, and reaching for it skips what actually lands. Most measure-triggered findings become a letter in a file, a quiet contraction of block time, a schedule rebuilt without discussion, a quality component in a compensation formula. Those carry no threshold, no appeal and nothing reportable, which makes them harder to argue with rather than easier. None of it is confined to surgeons: privileged advanced practice clinicians run the same evaluation at smaller volumes, which makes the arithmetic worse for them.
The reliability of the data used in that evaluation has never been published. I went looking for it. The rigor followed the money.
The objection
I should name my own interest, and the obvious version is not the important one. I lead the clinical side of an anesthesia company, so I am one of the parties that issues numbers like these. The sharper conflict is that an argument against holding individual clinicians to noisy numbers serves the person who employs clinicians. Discount accordingly. I am not arguing that no number should ever attach to a person, and that objection gets its due below.
The disclosure I actually owe is narrower than either of those. Two weeks ago I wrote that the system depends on someone in the room being willing to say how a number was built, and that this job has no audit and no owner. I was that person in this room. I said nuanced. And I have never had the longer conversation with every clinician these numbers reach. I am not certain I have had it completely with more than a handful. Ever.
The strongest objection is one a reader will find before finding me. Published work shows that adding cancer-specific variables to this kind of model did not change hospital quality rankings. At the hospital level that is frequently true. Volumes are large, rankings are coarse, and refining risk adjustment moves fewer institutions than people expect.
That is a finding about hospitals. The unit here is a person, whose volume is a fraction of an institution’s, and whose consequence is a privilege file rather than a payment adjustment. The paper is not wrong. It answered the question the system bothered to ask, which is the question with money attached to it.
The second objection is that if the measures are this soft we should stop measuring. That does not follow. Measurement is how a group finds a problem it cannot see from inside, and what the reliability finding argues for is a change of method rather than an abandonment: judge cases, not rates, when the rate cannot carry the judgment. When a person is on the other end of a number, the precision of that number is part of the number.
What is actually owed
Not that the number is soft. That sentence has no object, and the room hears a verdict on the whole enterprise. Say what it is soft about. This measure. This missing variable. This direction.
Name where it ran in their favor, because it does, and I have just spent a paragraph doing it months late. A leader who discloses softness only when it exculpates the institution has disclosed a preference rather than a fact.
Then pay the part that can be paid, and be specific about what that is, because a vague obligation is the same failure in a different suit.
The human read already exists on paper. Focused review is a human read. The mechanism is not missing — the trigger fires before anyone has asked whether the number could support it, and by then the clinician has been labeled. So condition the trigger. Before a ratio attaches to a named individual, someone writes down what the model adjusted for, what the denominator was, how that denominator compares to the case counts published work says a judgment about a person requires, and which direction the missing variables push. That document goes in the file next to the number.
None of that requires permission. Before you sign an evaluation packet or forward a scorecard to the person it describes, ask those four questions and write the answers next to the number. Your signature is yours. That sits inside someone’s authority in every organization I have worked in.
And say what still cannot be changed. The number will keep counting. Naming the limit is the difference between disclosure and the performance of disclosure.
What I offered that surgeon was sympathy. What was owed was arithmetic.
I do not know what would have happened if I had done that work and said it. The number would have counted anyway. The evaluation would have arrived on schedule. Whatever that ratio was going to do to that practice, it was going to do.
What would have changed is that the surgeon would have known which parts of it were real, and would not have had to take my word for the rest.
This is that conversation, held late and in the wrong room.
Past the Door is free, and it will stay free.
At some point I will open a paid tier. It will be for readers who want more than one piece a week, and the free posts will not get thinner when it happens. I am not there yet, and I would rather know now than guess.
If this is worth something to you, you can pledge a subscription. Substack holds it. You are not charged unless I turn payments on, and if I never do, nothing happens. What a pledge tells me is whether to build the paid tier at all.


