A health-risk score that predicted cost, not need

A widely used care-management algorithm predicted health costs as a proxy for illness, so Black patients were sicker than White patients at the same score.

Year
2019
Jurisdiction
United States
Sector
Healthcare: care management
Evidence base
Primary sources
Incident record
AIID 124
Harm
Discrimination in consequential decisions

What happened

Health systems use commercial prediction algorithms to pick patients with complex needs for extra care. Obermeyer and colleagues studied one widely used algorithm of this kind and found that, at a given risk score, Black patients were considerably sicker than White patients, as shown by signs of uncontrolled illness 1.

The bias arose because the algorithm predicts health-care costs rather than illness, and unequal access to care means less is spent on Black patients than on White patients. Remedying the disparity would raise the share of Black patients receiving additional help from 17.7% to 46.5% 1. The AI Incident Database records the case against the vendor, as reported 2.

Failure mode

Label choice. The target the model learned (future cost) stood in for the construct the programme cared about (future need), and the proxy was itself shaped by unequal access. Measured against its own label, the model could look accurate and still be wrong about who needed care 1.

Which control would have caught it

The model card is where the gap between label and construct is written down and owned: what the programme wants to predict, what the model actually predicts, and why the difference is acceptable. An eval gate that checks calibration by group against a measure closer to the construct (for example, active chronic conditions) shows whether equal scores mean equal need.

Patterns: Model Card as Control Evidence · Eval Gate in CI

The evidence that would have existed

What an auditor could have read, and the stack layer that produces it.

  • L2 Model card section on label validity: construct, proxy label and known gaps, signed by the clinical owner
  • L3 Eval report of calibration by group against a health measure, with the gate's verdict
  • L5 Change record whenever the label or the programme-entry threshold changes

Obligations it touches today

As of 2026-09-24. Mappings are illustrative, not a claim of conformity.

  • EU AI Act Art. 10 For high-risk systems, training data must be examined for possible biases likely to affect health and safety or fundamental rights 3.
  • EU AI Act Annex III, point 5(a) Evaluating eligibility for essential public assistance benefits and services, including healthcare services, is high-risk when done by or for public authorities 4. Whether a care-management score used by a private provider falls there, or under a medical-device route, is a legal classification question.

Sources

  1. [1] Dissecting racial bias in an algorithm used to manage the health of populations (Science 366(6464):447-453). Obermeyer, Z., Powers, B., Vogeli, C., Mullainathan, S.. 2019-10-25. https://doi.org/10.1126/science.aax2342 (verified: primary)
  2. [2] AI Incident Database, Incident 124: Optum Algorithmic Health Risk Scores Reportedly Underestimated Black Patients' Needs. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/124/ (verified: primary)
  3. [3] EU AI Act Art. 10 (data and data governance; examination of training data for possible biases). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_10 (verified: primary)
  4. [4] EU AI Act Annex III (high-risk uses; point 3(b) evaluating learning outcomes, 4(a) recruitment and selection, 5(a) eligibility for essential public assistance benefits and services). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#anx_III (verified: primary)