The caveat was true. The conclusion was not.
A trustee asked whether the pattern held in other patient populations. I had built the slide that said the question did not matter.
Some years ago I stood in front of a board quality committee and explained why a metric in our CMS star rating did not tell us much.
To be fair, we didn’t like the picture the number painted, and the explanation was sound. The measure was derived from fee-for-service Medicare patient encounters, which were under a fifth of our volume. Medicare Advantage already made up a larger share of our payer mix and was not included in the measure at all. Add Medicaid at roughly fifteen percent and a large commercial payer mix, and most of our encounters were invisible to the number on my slide. Whatever the measure showed about those patients, I said, it could not be read across the rest of the population. I had built the slide. I had written the assessment on it. Our management team had reviewed it in prep and agreed, which is how a position arrives in a committee already settled.
The gap I was pointing at was real, and real enough that CMS has since started closing it. Condition-specific readmission and mortality measures are being rebuilt to include Medicare Advantage patients, beginning with the FY 2027 program year. It was not closed then.
A trustee asked whether the findings held in other patient populations.
He was an attorney. He had no clinical training and he was not disputing the medicine. His question was whether the pattern in that population would show up in other populations if we measured them the same way, and if it did, whether the small number was a limitation or a sample.
I did not have an answer. I am not sure the question fully landed on me in the room, because I had walked in with the matter closed.
We took it to our analytics team afterward. It did not occur to me to ask the clinicians who had cared for those patients. By the time a measure like that reaches a board the cases are years back, and no one reconstructs a cohort from memory. What a clinician recalls is the case that stood out, which is a biased sample of exactly the wrong kind. They built a comparable assessment — same outcome, same logic, risk adjustment re-fit on our own data — that could run on any population, and we ran it on one.
Then, because I did not accept the first result, on another. Then across several payer and clinical groups, until I stopped asking for one more.
I would like to call that rigor. Running a test again because you dislike the answer is rigor only by accident, and I could not tell you today which of those runs was diligence and which was hoping.
The number that measure produced was specific to the population it was built for, exactly as I had said. The pattern it was pointing at showed up about as often in every group we looked at.
Then the rest of the team had to arrive at the same place, which took considerably longer than the analysis did. By that point the position on my slide was something we had been claiming for months or years, long before the board meeting where the question was asked. The finding ran against it.
I had been right about the number and wrong about what it entitled me to conclude. The caveat was accurate. The dismissal it was suggesting was not, and I was the person in that room best equipped to know the difference.
That is what I want to write about, because I do not think it was a lapse. I think it is the ordinary condition of the job.
A board does not fail to check the translation. The translation arrives pre-hardened, and the person who pushes on it is outnumbered by arrangement rather than by argument.
The room, and the room before it
The quality committee meets quarterly. An hour, sometimes ninety minutes, and quality is the second item because the first one is approving the last set of minutes. A lot of boards moved quality to the top of the agenda a decade ago, which was the right change and did not alter a line of what is in the report.
A dashboard goes up. Mortality index, PSI-90 (a composite of patient safety indicators), hospital-acquired conditions, a star rating, a Leapfrog grade. Most of it is green. The quality officer built the packet, the chief medical officer takes the clinical questions, the nursing officer takes the falls and the pressure injuries, and the committee accepts the report and moves on.
That hour is the mechanism by which a governing body discharges an obligation it is surveyed against. The Medicare Conditions of Participation put quality of care on the governing body by name.
But the hour is not where the interpretation gets decided. It gets decided in the prep session, among people with different training and the same incentive, and the committee receives the settled version. When the trustee spoke, he was not questioning a slide. He was breaking a room’s agreement, alone, without the vocabulary, against people who had already agreed.
That is the part I had never looked at squarely, and it is why the hour is not the problem.
What the number is made of
Here is what I know now that I would have said I already knew then.
The mortality index and PSI-90 on that screen are not built from the chart. They are built from the coded discharge abstract: the diagnosis codes, and the flag that says whether a condition was present on admission.
That sounds like a technical detail and it is the whole thing.
A mortality index is observed deaths over expected deaths, and the expected side is assembled from coded comorbidities. A more thorough coding effort raises what the hospital was expected to have and lowers the index. Present-on-admission is the cleanest version of the same effect: whether a pressure injury was flagged as present when the patient arrived decides whether it counts against you at all. One field. No change in care.
So the sophistication of a documentation program predicts measured quality independently of the care delivered. A hospital that invests there looks better without a single clinical practice changing.
That is not fraud and it is not gaming the system. Coders who capture comorbidities accurately are doing real and necessary work, and a record that reflects how sick the patient actually was is a better record. Incomplete coding distorts mortality indices and patient safety indicators downward, exactly as thorough coding lifts them. I have said to more physicians than I can count, document well, you should get credit for the acuity of the care you have already provided.
Adjusting a measure for clinical risk is widely accepted. Adjusting it for social risk is less so. Preventable hospitalizations are strongly associated with patients’ social and cognitive circumstances, and the physician value-based payment measures built on them do not capture that. The same omission runs through the hospital-side measures. A hospital serving a poorer and sicker population is measured as a worse hospital, partly for reasons that belong to the population.
The rating systems disagree with one another. Across 2,384 hospitals studied in 2023, CMS Care Compare and Leapfrog were discordant 70 percent of the time, severely discordant in a quarter of cases. Healthgrades and U.S. News disagreed on orthopedic procedure quality between 48 and 61 percent of the time. They are not measuring the same thing. Leapfrog is a safety composite. The star rating spans mortality, safety, readmissions, experience and timeliness on its own weighting. Neither grade is wrong, which is worse than one of them being wrong. Nothing on the dashboard says so.
Every one of those facts is true. Every one of them is also available as a reason to set a number aside, especially when the number tells a story you don’t like.
The problem with being the person who knows this
A methodological limitation is not an opinion. It is a fact about an instrument, and the person who can state it accurately is doing the room a service.
That same fact is the most efficient way to make an inconvenient number go away, and it is available only to the person with the vocabulary to state it. Nobody else in that committee could have made it stick the way I did. If a trustee had tried, I would have corrected him, and I would have been right on the facts while being wrong about the part that mattered.
That is the asymmetry, and I do not think there is a clean way out of it. The caveat and the dodge are the same sentence. They are told apart by what the speaker does next, and the speaker is the only witness to that.
Which is exactly why the check has to come from somewhere other than the person making the claim, and nothing in the structure requires one.
Evaluating the translation would require the clinical knowledge the translation exists to provide. There is usually a physician on the committee, sometimes chairing it, and they have almost never built or validated a measure. Presence is not a check. Every input to the number is audited hard: the abstraction, the coding, the DRG assignment, the quarterly chart re-abstraction on the measures that carry one, with a payment consequence attached. None of the audits touch what the number was said to mean. There is no audit of the translator.
There is also no literature on it that I could find. The translation has never been studied as a variable, which means what I am describing is built from the incentives and from the rooms I have been in. Weigh it that way.
There is one group that could check the clinical half of it. It is the staff who were there, and no structure carries their answer into the boardroom. By the time the number arrives, the answer may not exist anymore.
The incentives run the wrong direction on both sides of the seat, and the first one is written down. A share of what that executive is paid is tied to the measures on the screen, and the plan that ties it goes to a committee of the same board. Upward, naming the softness means telling a board that reports it has been governing on, several of which you delivered, were thinner than they sounded. Downward it is worse. Chartis found that barely half of physicians report significant trust that their executives are honest and transparent with them, which is a finding about executives before it is a constraint on them, and it is still the condition you speak into. Tell the clinical staff the dashboard is soft and you have handed the most cynical version of the room their loudest argument. They will not repeat the instrument has known error. They will repeat the numbers are theater, and the quality program loses credibility.
There is a venue that sounds like it was built for the harder conversation. Every board quality committee I have sat with holds an executive session, and in every one of them the chief executive and the chief operating officer were in the room alongside the clinical officers. It was not a session without management. It was the same room with the door closed, and it mostly got spent on credentialing and personnel.
So the pressures converge on one behavior. Present the report. Answer what you are asked. Do not volunteer the architecture. And the system contains no penalty for that choice. The green dashboard is the state in which the meeting ends on time and no one asks a hard question.
This is a decision with no owner. The report has an owner, with a department and a production calendar. What has no owner is the interpretation on top of it.
I should share my own standing. I lead the clinical side of an anesthesia company, and hospitals pay us for clinical leadership, which means I am paid to be right about the argument I am making. Discount accordingly. The room I own now is a client facility’s quality committee rather than a health system board’s, and the asymmetry there is identical and smaller.
What the trustee actually did
He did not know more medicine than the people in front of him. He knew his own job, which is testing whether a claim generalizes past the case it was built on, and he applied it to material that was not his.
He also did the one thing the arrangement was not built to absorb: he asked a question that could be answered. Not are you sure, which invites reassurance from the expert and ends there. Would the pattern show up in another population measured the same way — which is checkable, which could not be settled in the room, and which therefore had to leave it.
It indicts the prep session rather than the meeting. The instrument that eventually answered him did not exist because it had never been needed. It took a few weeks to build. It could have been built before the slide, and the reason it was not is that the position on the slide had stopped being a finding and become a thing we said.
The more insidious failure is the one I would watch for. A conclusion gets examined once, at the moment it is formed. After that it gets repeated, and repetition is not review.
There is a reason it runs that way, and it is not laziness. Nothing on a governance calendar schedules a second look at a position already taken. Packets are built forward, and last quarter’s interpretation is this quarter’s starting point.
Underneath that, healthcare leaders have had more data than they can use since the EMR arrived, and most of us still decide by instinct, because instinct has been working well enough. A career of adequate outcomes does not send anyone back to examine the belief underneath them.
What I carried into that committee was an instinct with a methodology attached to it. The limitation was real. It was also the reason I did not have to think about the number again.
Trusting your gut is easier than challenging your beliefs.
What the obligation actually is
It is not better numbers. The numbers will stay imperfect.
The obligation is to change what the room is capable of asking, and then to answer it when it asks.
Report the measure and your confidence in the measure. CMS publishes interval estimates on its mortality and readmission measures and uses them to say whether a hospital differs from the national rate. Control limits do the same on a trend. A number without an error bar is a claim, not a finding.
Name which measures are documentation-sensitive, and which move with your payer mix, before the board reads a rise as clinical improvement rather than after.
Bring one case per meeting the dashboard did not catch, say who caught it, and tell them what the committee did with it. Not as color. As evidence about the instrument. De-identified, or in the session where clinical case review belongs. The patient did not consent to be governance evidence.
Say it out loud when a metric moved because coding changed. Every time. That is the sentence that costs the most and buys the most, and it is one I have swallowed in the moment.
And when you reach for a limitation, say what would have to be true for you to be wrong, before anyone asks. That is the one I did not do. The limitation I named was real, I never stated what evidence would change my mind, and so an accurate sentence did the same as a closed door.
The test has a second half, and I failed that one too. Count how many confirmations you require when the answer is unwelcome, and compare it to how many you required when the answer was the one you brought in. I ran that assessment on population after population. I had run the original on one.
None of that requires a board seat or a title. The same function runs at every level where a number goes up a rung: the service line director signing the dashboard that feeds the packet, the department chair presenting to the medical executive committee, the medical director reviewing quality data with a group. Smaller room, same two pressures, same sentence available.
The objection
The obvious one is that I have written a piece concluding that the person the board most needs has my job, at a company paid to supply it.
Fair. What I can put against it is weaker than I would like. The citation physician executives keep in a back pocket is Goodall, which reported roughly 25 percent higher quality scores at physician-led hospitals among top-ranked institutions. A 2022 analysis in JAMA Network Open went wider, drawing on an American Hospital Association file of 6,162 hospitals, roughly a third of which carried HCAHPS ratings and Leapfrog grades. In multivariable models it found no significant association between a physician chief executive and either HCAHPS ratings or Leapfrog grades. The one thing that survived was patients’ willingness to recommend the hospital.
So the strongest evidence for the claim I have an interest in believing is a finding a much larger sample did not reproduce. It failed in exactly the way my slide did, which I notice. I would rather say it than have you tell me later.
What survives sits at the right level and is not as strong as I hoped. Hospitals whose boards had a quality committee at all — about six in ten when it was surveyed in 2006 — showed lower mortality across six common conditions than those that did not. That is a finding about whether the committee exists. It says nothing about who sits on it, or what they caught.
For the claim I wanted to make for this article — that clinical judgment on that committee is what catches a bad translation — my notes had a citation. A 2015 finding, in a quality and safety journal, from a researcher whose work on hospital governance I have cited for years: hospitals whose board quality committees included members with clinical expertise did better on process and outcomes. Author, journal, year, reference number. Every surface I would normally check.
It got opened this week, because somebody noticed in reviewing this piece that my own notes described the same reference two different ways. It is a different paper by different authors in a different journal, and it is a qualitative study built on ten interviews in the English National Health Service — a system whose board structure and accountability chain are not the ones I have spent this essay describing. Good work, and not evidence for anything I wanted to say. The closest honest summary I can give you is an OECD scan across nineteen countries: the evidence that doctor involvement in hospital governance improves performance is limited.
I am sharing that example consciously rather than pretending it didn’t happen. The citation was correct at every surface I did check and wrong underneath, which is this entire essay arriving in my own footnotes. And I did not catch it. It surfaced the way the number did. Somebody noticed something that did not line up, and asked.
In any case the story above does not argue for putting a clinician in the room. There was one. I built the slide.
It argues for putting someone in the room whose job is to ask whether the claim generalizes, and for that person to be hard to ignore.
The question
He was more right than I was, and he could not have told you why in clinical terms. He did not need to.
If you sit in a seat like mine, or you already run a smaller version of it, you know which caveat you did not raise last quarter. And if you chose the bedside, you know which note you wrote quickly on a night you were carrying too much, and the number was built out of what a coder made of that.
The question I would put to either of you is the one I could not answer in that room. What would it take to show you that you are wrong — and have you built the thing that could?


