Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when evaluating the reliability (score consistency/stability) of an AI/ML selection tool — Concern 6 of Tippins, Oswald & McPhail (2021). Covers reliability as an absolute requirement, what stability means for AI scores, evidence that machine scoring can be as or more reliable than human scoring, the questionable reliability of facial-emotion analysis (including across skin tone, disability, and altered features), and confounds from individual differences in the data generated (e.g., extrave
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-02 | ✓→✓ | = Same ✓ | -14% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 25% | 0% |
Like validity, reliability is an absolute requirement for any test. The Uniform Guidelines require reliability to be evaluated and reported (§§14C(5), 15B(7), 15B(8)); the Principles (p. 22) require predictor scores to exhibit adequate reliability and to identify the conditions of measurement across which one wishes to generalize.
For how to estimate reliability, use criterion-related-validation (predictor/criterion reliability) and ai-model-outputs-audit.
To be reliable, scores must measure KSAOs that are relatively stable and consistent over time and setting:
(memorization notwithstanding).
equipment unless those are job relevant.
skill) or reduced irrelevancy (less anxiety, better understanding of the protocol) — not artifacts.
normally required (e.g., multikey or mouse facility).
Reliability evidence for technologically enhanced measures is often minimal or difficult to obtain; when available, results are mixed.
interviewer-scored — with substantial agreement with well-trained humans. Because human scorers in operational settings are notoriously prone to error and low interrater agreement, there's long-recognized potential for greater reliability and fairness in algorithmic scoring (Kuncel et al., 2013) so long as job-relevant information is what's being scored.
expressions can reliably signal categories of emotion in human judgment (Cowen & Keltner, 2019), studies suggest the faces of people with darker skin are harder to evaluate via AI (Singer & Metz, 2019). It is unclear how the tool treats people with injuries, disabilities, medication effects, or altered features (scars, tattoos) — a reliability and fairness problem.
Threats to reliability can come from individual differences in the data collected, not just how they're measured. For example, more extraverted applicants provide more detail in written or spoken responses, giving the algorithm more key words that may relate to other traits (cognitive ability, verbal skills) not necessarily related to extraversion. So big data, ML, or both can end up affecting the reliability of other traits the tool claims to measure — a subtle source of construct-irrelevant variance.
selection?
enhanced assessments?
ai-validity-evidence · ai-ml-methodology-evaluation · ai-dynamic-models-and-revalidation (stability vs. frequent updating) · ai-candidate-data-control (disability/appearance fairness) · criterion-related-validation (reliability estimation) · fairness-and-bias-analysis · ai-model-outputs-audit
Source: Tippins, Oswald & McPhail (2021), Concern: "Reliability."
Other measured skills in the registry, with their headline benchmark lift.