Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when evaluating fairness and bias of a selection procedure — distinguishing the several meanings of "fairness," testing for predictive bias (differential prediction via moderated regression), and examining measurement bias (DIF, item sensitivity review). Covers what subgroup-mean differences do and don't imply, when bias analyses are warranted, and the statistical pitfalls. Triggers: "adverse impact vs bias", "differential prediction", "predictive bias", "measurement bias", "DIF analysis", "
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -4% | 0% |
Two distinct ideas that are routinely conflated. Keep them separate.
"Fairness" has no single agreed definition (statistical, psychometric, or social). Recognized meanings include:
definition of fairness: group outcome differences alone do not indicate bias — though they should trigger heightened scrutiny for possible bias.
feedback, retest opportunities, reasonable accommodation, mode of administration).
standing without being advantaged/disadvantaged by construct-irrelevant characteristics (age, race, ethnicity, gender, SES, cultural/linguistic background, disability).
There is broad agreement that equitable treatment, access, bias, and scrutiny when subgroup differences appear are important — but no agreement that "fairness" can be uniquely defined in terms of any one of them.
Bias = systematic error that differentially affects the performance of different subgroups.
groups (a predictor–criterion relationship issue).
scores for a subgroup (a score issue, for predictors or criteria).
A subgroup-mean difference (adverse impact) is a negative consequence, but it is evidence against validity only if it traces to a measurement property of the procedure (i.e., bias). If the group difference on the procedure mirrors a real difference in the work-relevant outcome (i.e., no predictive bias), the consequence is a policy issue for the user, not a validity defect.
Test via moderated multiple regression (MMR): regress the criterion on the predictor, subgroup membership, and their interaction. Slope and/or intercept differences signal predictive bias. MMR is preferred over comparing separate subgroup correlation coefficients.
underprediction signals bias against that group. Simply knowing slopes/intercepts differ doesn't answer it. (In U.S. cognitive-ability research, slope differences are rare; when intercept differences occur they typically take the form of overprediction of minority performance — Schmidt, Pearlman, & Hunter, 1980; and corrected analyses, e.g., Berry & Zhao, 2015, still find little underprediction.)
Sackett, 2017).
composite, not each test separately).
range restriction, and predictor unreliability all reduce power to detect slope/intercept differences.
(not observed parameters).
Predictive bias and mean differences can exist independently; analyze predictive bias when there's compelling reason to question whether predictor and criterion relate comparably across subgroups and appropriate data exist. Where relevant research exists, generalized evidence can inform the question.
Construct-irrelevant variance raising/lowering scores for a subgroup — hard to detect because it requires comparing an observed score to a true score. Approaches:
scorers) for language/content that could carry differing meaning across subgroups or be demeaning/offensive. Value depends on content; use is a matter of professional judgment.
with the same total score (or same IRT true score) perform differently. Notes:
for cognitive tests it's common to find roughly equal numbers of items favoring each subgroup, netting to little test-level bias.
exist. Especially useful in cross-cultural / linguistically different testing.
criterion-related-validation · selection-decisions-and-scoring (composites & subgroup tradeoffs) · candidate-accommodations (equitable treatment/access) · internal-structure-validation · technical-validation-report
Source: Principles (5th ed., 2018), "Fairness and Bias."
Other measured skills in the registry, with their headline benchmark lift.