Pick the fairness metric before the results.

Which fairness metric fits depends on the harm, on whether a trustworthy label exists, on which error costs more and on the legal frame. Answer those once, before the eval runs, and get the metric families to use, the checks beside them and the chapter's warnings.

Indicative, not legal advice and not a conformity claim. Nothing you enter leaves your browser.

JavaScript is off or has not loaded, so the live result and the exports are not available. The questions, the criteria and the guide below still work as a worksheet you can read and fill in by hand.

Built from 16. Fairness and explainability of the Body of Knowledge. Runs entirely in this page: no account and no upload.

It names the exports and the link.

What does the system do to people?

The metric follows the harm, and the harm follows the use case. Chapter 16: Choosing a fairness metric by use case.

Is there a trustworthy ground-truth label?

Error-rate metrics compare predictions with the true outcome. A label that stands in for the real target passes every accuracy test and still carries measurement bias. Chapter 16: Where bias enters the lifecycle.

Which error costs more?

Take it from the use-case record's error appetite, not from which metric the model passes. Only for allocation harms. Chapter 16: Choosing a fairness metric by use case.

Which legal frame governs the decision?

The law decides which disparity is unlawful; the engineer builds the measurement so Legal has something true to decide on. Chapter 16: Disparate treatment and disparate impact.

Can you use the protected attribute at evaluation time?

Both proxy tests and every group metric need it; that is the privacy tension every fairness programme meets. Chapter 16: Protected characteristics, proxies and the data you need to test.

How the tree reads chapter 16

The metric follows the harm, and the harm follows the use case [1]. Allocation harms (a system extends or withholds an opportunity, a resource or information) and quality-of-service harms (it works less well for some people) need different metrics, and stereotyping is a third kind [2]. Add the cost of each error type and the legal frame, and the choice narrows. Because the metrics conflict when base rates differ [4][5], the choice is a governance decision with an owner, taken before the results are seen.

The tree, by hand

  1. Allocation, with a trustworthy label. Take the metric of the costliest error: missing a qualified person, equal opportunity; a wrongful cut or flag, false-positive-rate parity; both, equalised odds [3]; a positive decision whose value depends on being right, predictive parity; a score read as a probability, calibration within groups.
  2. Allocation, with a label that may be a proxy. Review the label first (where bias enters the lifecycle); meanwhile measure selection rates, which need no label.
  3. Allocation, with no label yet. Selection rates and the adverse-impact ratio, the counterfactual flip test, and the error-based metric once outcomes arrive.
  4. Quality of service. The worst-group error rate, with intersectional error.
  5. Generative output. The counterfactual flip rate and a per-group quality floor, with stereotype probes and refusal-rate gaps.
  6. Then the legal frame. US employment adds the selection-rate AIR and the four-fifths screen [6][7] and rules out group-specific cut-offs [8]; credit pairs calibration with the approval-rate AIR; benefits put wrongful cuts first; in the EU, high-risk systems examine data for biases and may use special-category data only under Art. 4a's conditions [9].

The metric families

What each family equalises, when it fits and what to watch
Metric Holds when Fits when Watch for
Demographic parity and the adverse-impact ratio Selection rates are equal across groups; the adverse-impact ratio (AIR) is its ratio form. The opportunity should be shared regardless of measured outcome; it needs no outcome label. Ignores different base rates, and can be met by selecting unqualified members of a group. The four-fifths rule is a trigger for investigation, not a pass mark.
Equal opportunity True-positive rates are equal across groups. Missing a qualified person is the main harm (hiring, admissions, access to care). Leaves false positives unconstrained.
Equalised odds True-positive and false-positive rates are both equal across groups. Both errors are costly. Harder to satisfy; may cost accuracy for all groups.
False-positive-rate parity False-positive rates are equal across groups. Wrongly cutting or reclaiming, the costliest error in benefits eligibility and recovery. Leaves missed eligible people unconstrained; pair it with appeal outcomes by group.
Predictive parity Precision, the share of positive decisions that are right, is equal across groups. A positive decision triggers action whose value depends on being right (fraud referral). Incompatible with equal error rates when base rates differ.
Calibration within groups Among people scored s, a fraction s are positive, in every group. Scores are consumed as probabilities (credit pricing, clinical risk). A calibrated score can still produce very different error rates.
Worst-group error rate The error rate of the worst-served group, reported beside the average. Quality-of-service harms: the system fails for a group of users (speech, vision, document search). Needs a labelled evaluation set drawn from the people actually served.
Counterfactual flip rate Change only the protected attribute or its textual markers, hold everything else fixed, and measure how often the decision or the generated text changes. The most practical fairness eval for LLM-based systems: group labels for outputs rarely exist, while paired prompts are easy to generate. An approximation of counterfactual fairness, which needs a causal model to compute exactly.
Per-group quality floor Output quality stays above a set floor for every group. Degraded or demeaning output for a group is the costliest error of a generative assistant. Set the floor, and the rating method, before the results are seen.

The chapter's use-case table

From choosing a fairness metric by use case: a starting point, not a rule. What makes a choice defensible is that it is written down with its reasons before the eval runs and reviewed by someone who represents the affected people.

Use case, costliest error and metrics
Use case Harm type Costliest error Primary metric Secondary checks Legal frame
CV screening, promotion Allocation Rejecting a qualified candidate Selection-rate AIR; equal opportunity Intersectional AIR; proxy scan Title VII, Uniform Guidelines, NYC LL144; AI Act Annex III point 4
Credit approval and pricing Allocation Both: wrongful denial and unaffordable credit Calibration within groups; approval-rate AIR Error-rate gaps; reason-code consistency ECOA and Regulation B, FCRA; AI Act Annex III point 5(b)
Benefits eligibility and recovery Allocation (punitive when reclaiming) Wrongly cutting or reclaiming a benefit False-positive-rate parity Predictive parity; appeal outcomes by group Equality law; GDPR Art. 22; AI Act Annex III point 5(a)
Clinical triage Allocation (need-based) Missing a person in need Equal opportunity; calibration Label-validity review (cost versus need) Medical-device and equality law
Speech, vision, document search Quality of service Failing for a group of users Worst-group error rate Intersectional error Accessibility and equality law
Generative assistant Quality of service; stereotyping Degraded or demeaning output for a group Counterfactual flip rate; per-group quality floor Stereotype probes; refusal-rate gaps Equality and consumer law

What this is not

Not legal advice and not a conformity claim. The law decides which disparity is unlawful and which fix is allowed; the tree only restates chapter 16, which names the toolkits that compute these metrics as examples of a category, not endorsements. Record the choice, the threshold and the approver in a Policy Card, and let the eval gate read it.

Terms used here

Sources

  1. [1] 16. Fairness and explainability for practitioners: "Choosing a fairness metric by use case", "Group fairness metrics", "The impossibility results", "Intersectional and subgroup testing", "Where bias enters the lifecycle" and "Protected characteristics, proxies and the data you need to test", the text this tool restates. AI Governance Engineering Body of Knowledge. 2026-09-24. https://aigovernanceengineer.com/bok/fairness-and-explainability (verified: primary)
  2. [2] Fairlearn user guide, "Fairness in machine learning" (allocation, quality-of-service and stereotyping harms; disparity metrics as ratios or differences). Fairlearn project. 2026. https://fairlearn.org/main/user_guide/fairness_in_machine_learning.html (verified: primary)
  3. [3] "Equality of Opportunity in Supervised Learning" (M. Hardt, E. Price, N. Srebro; equalised odds and equal opportunity). arXiv 1610.02413. 2016-10-07. https://arxiv.org/abs/1610.02413 (verified: primary)
  4. [4] "Fair prediction with disparate impact: A study of bias in recidivism prediction instruments" (A. Chouldechova; predictive parity and equal error rates cannot both hold when prevalence differs). arXiv 1703.00056. 2017-02-28. https://arxiv.org/abs/1703.00056 (verified: primary)
  5. [5] "Inherent Trade-Offs in the Fair Determination of Risk Scores" (J. Kleinberg, S. Mullainathan, M. Raghavan; calibration within groups and balance for the positive and negative class cannot hold together except in highly constrained special cases). arXiv 1609.05807. 2016-09-19. https://arxiv.org/abs/1609.05807 (verified: primary)
  6. [6] 29 CFR 1607.4(D), Uniform Guidelines on Employee Selection Procedures (1978): adverse impact and the "four-fifths rule", with the statistical and practical significance and small-numbers caveats. eCFR (text as of 2026-09-01). 2026-09-01. https://www.ecfr.gov/current/title-29/subtitle-B/chapter-XIV/part-1607/section-1607.4 (verified: primary)
  7. [7] Automated Employment Decision Tools: Frequently Asked Questions (Local Law 144 of 2021: independent bias audit within the past year; impact ratios across sex, race/ethnicity and intersectional categories; published summary; no specific action required). NYC Department of Consumer and Worker Protection. 2023-06-29. https://www.nyc.gov/assets/dca/downloads/pdf/about/DCWP-AEDT-FAQ.pdf (verified: primary)
  8. [8] 42 U.S.C. § 2000e-2(k) and (l) (disparate impact burden of proof; no adjusted scores or different cut-off scores by race, colour, religion, sex or national origin). Legal Information Institute, Cornell Law School. 2026. https://www.law.cornell.edu/uscode/text/42/2000e-2 (verified: secondary)
  9. [9] Regulation (EU) 2024/1689 (AI Act), consolidated text as amended by Reg. (EU) 2026/1744: Art. 4a (special categories of personal data for bias detection and correction), Art. 10(2)(f)-(g) (examination for and mitigation of biases), Annex III points 4, 5(a) and 5(b). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_10 (verified: primary)