A health-risk score that predicted cost, not need
A widely used care-management algorithm predicted health costs as a proxy for illness, so Black patients were sicker than White patients at the same score.
One incident read against the controls of AI governance and its frameworks.
- Year
- 2019
- Jurisdiction
- United States
- Sector
- Healthcare: care management
- Evidence base
- Primary sources
- Incident record
- AIID 124
- Harm
- Discrimination in consequential decisions
What happened
Health systems use commercial prediction algorithms to pick patients with complex needs for extra care. Obermeyer and colleagues studied one widely used algorithm of this kind and found that, at a given risk score, Black patients were considerably sicker than White patients, as shown by signs of uncontrolled illness 1.
The bias arose because the algorithm predicts health-care costs rather than illness, and unequal access to care means less is spent on Black patients than on White patients. Remedying the disparity would raise the share of Black patients receiving additional help from 17.7% to 46.5% 1. The AI Incident Database records the case against the vendor, as reported 2.
Failure mode
Label choice. The target the model learned (future cost) stood in for the construct the programme cared about (future need), and the proxy was itself shaped by unequal access. Measured against its own label, the model could look accurate and still be wrong about who needed care 1.
Which control would have caught it
The model card is where the gap between label and construct is written down and owned: what the programme wants to predict, what the model actually predicts, and why the difference is acceptable. An eval gate that checks calibration by group against a measure closer to the construct (for example, active chronic conditions) shows whether equal scores mean equal need.
Patterns: Model Card as Control Evidence · Eval Gate in CI
The evidence that would have existed
What an auditor could have read, and the stack layer that produces it.
- L2 Model card section on label validity: construct, proxy label and known gaps, signed by the clinical owner
- L3 Eval report of calibration by group against a health measure, with the gate's verdict
- L5 Change record whenever the label or the programme-entry threshold changes
Obligations it touches today
As of 2026-09-24. Mappings are illustrative, not a claim of conformity.
- EU AI Act Art. 10 For high-risk systems, training data must be examined for possible biases likely to affect health and safety or fundamental rights 3.
- EU AI Act Annex III, point 5(a) Evaluating eligibility for essential public assistance benefits and services, including healthcare services, is high-risk when done by or for public authorities 4. Whether a care-management score used by a private provider falls there, or under a medical-device route, is a legal classification question.
Sources
- [1] Dissecting racial bias in an algorithm used to manage the health of populations (Science 366(6464):447-453). Obermeyer, Z., Powers, B., Vogeli, C., Mullainathan, S.. 2019-10-25. https://doi.org/10.1126/science.aax2342 (verified: primary)
- [2] AI Incident Database, Incident 124: Optum Algorithmic Health Risk Scores Reportedly Underestimated Black Patients' Needs. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/124/ (verified: primary)
- [3] EU AI Act Art. 10 (data and data governance; examination of training data for possible biases). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_10 (verified: primary)
- [4] EU AI Act Annex III (high-risk uses; point 3(b) evaluating learning outcomes, 4(a) recruitment and selection, 5(a) eligibility for essential public assistance benefits and services). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#anx_III (verified: primary)