Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when auditing the scores an AI/ML personnel assessment produces — Component 6 of the Landers & Behrend (2023) framework. Covers evaluating the quality of model predictions: reliability (consistency over time and repeated administrations), validity evidence (do scores reflect the claimed constructs and predict the outcome), appropriateness of the cross-validation given generalizability claims, and subgroup differences across protected classes and their intersections. Triggers: "evaluate AI as
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 61% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 36% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 48% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 57% | 0% |
The payoff component: regardless of how the model was built, are its scores good? This is where the AI audit reconnects most directly to classic psychometrics and the SIOP Principles — apply the personnel-selection skills here, on the model's actual operational outputs.
What it is: evaluation of the quality of the predictions the model generates.
Questions to ask: How was prediction quality evaluated — e.g., for psychometric reliability and validity? How was cross-validation conducted, and was it appropriate given the claims about model generalizability?
Apply it (focal example):
who records the same interview twice, or whose video is scored repeatedly, should get stable scores.
relate to the intended outcome? Don't accept a predictive R² as construct validity.
color, national origin, religion, disability, age — and combinations of classes (intersectional subgroups)? AI's high-dimensional inputs make intersectional effects both more likely and easier to overlook.
administrations, alternate prompts, re-scoring) and estimate the matching reliability. For stochastic models or pipelines with random elements, test score stability under re-run. See criterion-related-validation (reliability of measures).
construct evidence that scores predict the defensibly defined outcome (tie back to the Component-2 criterion audit). Beware overfitting: a cross-validated estimate beats holdout, which beats temporal — confirm the cross-validation matches the generalizability claim (a tool used on future applicants needs evidence that survives temporal validation, not just k-fold). See criterion-related-validation and internal-structure-validation.
difference (adverse impact) is a scrutiny trigger, not a verdict of bias (Lens 2/3).
(the composite the model outputs), not internal sub-features; ask whether a subgroup is underpredicted, watch power, range restriction, and error-variance homogeneity. See fairness-and-bias-analysis.
remember psychometric bias may be expected and "fair" when groups truly differ on the construct (ai-fairness-lenses, Lens 3).
report what couldn't be analyzed.
population where the tool is deployed (the incumbent-vs-applicant / range-restriction problem from Component 1 resurfaces in the output evidence).
The focal developer claims the algorithm "predicts job performance equally well for all groups using appropriate modeling techniques." That claim is evaluated here — and the paper calls it questionable and in need of audit. Translate it into testable pieces: (a) equal predictive accuracy across groups, (b) no problematic differential prediction, (c) appropriate modeling — and require evidence for each.
ai-input-data-and-design-audit · ai-model-development-audit · ai-fairness-lenses · criterion-related-validation · internal-structure-validation · fairness-and-bias-analysis · selection-decisions-and-scoring
Source: Landers & Behrend (2023), Table 1 (Component 6).
Other measured skills in the registry, with their headline benchmark lift.