The property that a model's confidence matches its accuracy: of the cases scored 0.9, about nine in ten are right. Modern neural networks are often poorly calibrated 1, so calibration is measured in the eval gate per version and subgroup before any threshold is trusted.
- Developed in
- ch. 11, Calibration before thresholds
- Chapters
- ch. 11, AI Defined
- Contrast with
- Calibration within groups
- Source
- 1 numbered reference, listed below
Where it is used
5 chapters of the Body of Knowledge use the term. Each link opens the first section that does.
- 10 · Reading List Canonical papers: measurement, fairness and evaluation 2 mentions
- 11 · AI Defined Kinds of AI that change the governance problem 11 mentions
- 14 · Development Testing and validation 2 mentions
- 15 · Deployment Model types and deployment options 2 mentions
- 16 · Fairness & XAI Group fairness metrics 5 mentions
Patterns that use this term
3 pattern pages use the term, most mentions first.
- Use-Case Intake & Risk Tiering 1 mention
- Fairness Eval Suite 1 mention
- Drift & Fairness Monitor 1 mention
Related terms
Sources
- [1] On Calibration of Modern Neural Networks (modern networks poorly calibrated; ICML 2017; arXiv 1706.04599). Guo, Pleiss, Sun and Weinberger. 2017-06-14. https://arxiv.org/abs/1706.04599 (verified: primary)
Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .