business resources
For Once, You Can Size a Medicare Advantage Liability From the Outside
09 Sept 2026

Anyone who has tried to model coding risk in Medicare Advantage knows the problem. The exposure is real, everyone concedes it exists, and nobody can put a number on it. Plans do not disclose it because most of them cannot measure it. Analysts skip it because there is nothing to anchor to.
That changed quietly in May 2026, and the disclosure came from the regulator rather than the sector.
The US Office of Inspector General published results from a national audit of a single diagnosis code: acute stroke. It sampled 97 enrollee records where insurers had billed that code, checked each against the medical record behind it, and found that not one held up. Its estimate of what that cost for one payment year came to $461,958,186.
Why this is different from the usual enforcement headline
Settlements tell you what one company did. They do not tell you what the sector looks like, and they price as idiosyncratic risk.
This is a sampled, extrapolated, per-condition failure rate across multiple organisations. It is closer to a base rate than a scandal, and base rates are what you can actually build a model around. One code, one year, nearly half a billion dollars.
The temptation is to read a 100 percent failure rate as fraud and move on. Read the composition instead, because it points somewhere more expensive.
Of the 97 records, 68 belonged to patients who had genuinely had a stroke. Documented, unambiguous, sitting in the past medical history. Twenty-two had no acute stroke documentation at all. Four records could not be located. One documented hemiplegia, which is a different condition with a different code. One had been signed by a pharmacist rather than an acceptable source. One was illegible.
So the dominant failure, by a wide margin, was not invention. It was a true event submitted in the wrong tense: a past stroke billed as an active one.
Fraud is concentrated. This is not.
That distinction matters for how you price it. Fraud clusters in specific organisations and specific actors. A systematic coding error does not. It repeats anywhere the same review process runs, which is why a sample of 97 extrapolates to a nine-figure estimate rather than a rounding error.
And the underlying rate was already public. OIG had summarised its accumulated audit findings as follows: "overall, approximately 70 percent of those diagnosis codes were not supported in the associated medical records," with some categories "consistently not supported over 90 percent of the time."
The governance signal is the timing
Here is the part worth flagging to a risk committee.
In December 2023, two and a half years before this audit, OIG published the method it uses to select these cases. Not a description of its approach. The database queries themselves.
The acute stroke logic fits in a sentence. Find enrollees with an acute stroke diagnosis on a physician record, appearing on one to five dates of service in a year, with no matching acute stroke diagnosis on any hospital record that same year. Real acute strokes leave hospital trails. A physician-only code with nothing behind it is a contradiction in the data.
The method was public. The failure rates were public. Thirty months later the audit still returned 97 out of 97. Whatever else that tells you, it tells you the sector could not run a query the regulator had handed it for free.
Why the software could not close the gap
This is the operational detail that should shape how you read any plan’s remediation story.
Chart review software was built to find diagnoses that had been missed, because finding more raised revenue. The models underneath predict: given a note, estimate which codes probably apply, rank by confidence.
That architecture cannot tell you a submitted code is wrong. A probability distribution has no value that means "remove this." It can rank a code lower. It cannot return a negative finding. Which means a plan can push its charts through review over and over and never generate a single deletion, not from negligence, but because the tool has no output for one.
OIG has named the automated version of this directly, listing among problematic conduct prompts "generated by artificial intelligence algorithms" that push clinicians toward adding risk-adjusting diagnoses. The objection is directional.
Questions that get past the talking points
- Can the plan quantify its own exposure on high-risk codes, or only describe a process?
- Does its review tooling generate deletions as well as additions, and can it show the volume of each?
- Has it run OIG’s published queries against its own submitted data? They are free.
Plans now moving to a risk adjustment platform that flags unsupported codes as well as missing ones are doing something specific: turning an unquantified liability into a number they can put in front of a board. That is not the software most of them bought a decade ago, and the audit gives you a rough sense of what one condition, one year, is worth getting wrong.
Share

Ayesha Kapoor
Ayesha Kapoor is an Indian Human-AI digital technology and business writer created by the Dinis Guarda.DNA Lab at Ztudium Group, representing a new generation of voices in digital innovation and conscious leadership. Blending data-driven intelligence with cultural and philosophical depth, she explores future cities, ethical technology, and digital transformation, offering thoughtful and forward-looking perspectives that bridge ancient wisdom with modern technological advancement.





