Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when auditing the scores an AI/ML personnel assessment produces — Component 6 of the Landers & Behrend (2023) framework. Covers evaluating the quality of model predictions: reliability (consistency over time and repeated administrations), validity evidence (do scores reflect the claimed constructs and predict the outcome), appropriateness of the cross-validation given generalizability claims, and subgroup differences across protected classes and their intersections. Triggers: "evaluate AI as
.claude/skills/openmatter-network-ai-model-outputs-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 61% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 36% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 48% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 57% | 0% |
The payoff component: regardless of how the model was built, are its scores good? This is where the AI audit reconnects most directly to classic psychometrics and the SIOP Principles — apply the personnel-selection skills here, on the model's actual operational outputs.
What it is: evaluation of the quality of the predictions the model generates.
Questions to ask: How was prediction quality evaluated — e.g., for psychometric reliability and validity? How was cross-validation conducted, and was it appropriate given the claims about model generalizability?
Apply it (focal example):
who records the same interview twice, or whose video is scored repeatedly, should get stable scores.
relate to the intended outcome? Don't accept a predictive R² as construct validity.
color, national origin, religion, disability, age — and combinations of classes (intersectional subgroups)? AI's high-dimensional inputs make intersectional effects both more likely and easier to overlook.
administrations, alternate prompts, re-scoring) and estimate the matching reliability. For stochastic models or pipelines with random elements, test score stability under re-run. See criterion-related-validation (reliability of measures).
construct evidence that scores predict the defensibly defined outcome (tie back to the Component-2 criterion audit). Beware overfitting: a cross-validated estimate beats holdout, which beats temporal — confirm the cross-validation matches the generalizability claim (a tool used on future applicants needs evidence that survives temporal validation, not just k-fold). See criterion-related-validation and internal-structure-validation.
difference (adverse impact) is a scrutiny trigger, not a verdict of bias (Lens 2/3).
(the composite the model outputs), not internal sub-features; ask whether a subgroup is underpredicted, watch power, range restriction, and error-variance homogeneity. See fairness-and-bias-analysis.
remember psychometric bias may be expected and "fair" when groups truly differ on the construct (ai-fairness-lenses, Lens 3).
report what couldn't be analyzed.
population where the tool is deployed (the incumbent-vs-applicant / range-restriction problem from Component 1 resurfaces in the output evidence).
The focal developer claims the algorithm "predicts job performance equally well for all groups using appropriate modeling techniques." That claim is evaluated here — and the paper calls it questionable and in need of audit. Translate it into testable pieces: (a) equal predictive accuracy across groups, (b) no problematic differential prediction, (c) appropriate modeling — and require evidence for each.
ai-input-data-and-design-audit · ai-model-development-audit · ai-fairness-lenses · criterion-related-validation · internal-structure-validation · fairness-and-bias-analysis · selection-decisions-and-scoring
Source: Landers & Behrend (2023), Table 1 (Component 6).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 34,195 | 33,490 | -2% | 1 | 1 | 0% | 6,039 | 7,176 | +19% | 0 | 0 | — |
case-02 | fail→fail | 32,663 | 35,016 | +7% | 1 | 1 | 0% | 6,232 | 7,250 | +16% | 0 | 0 | — |
case-03 | pass→pass | 18,028 | 25,404 | +41% | 1 | 1 | 0% | 3,059 | 4,934 | +61% | 0 | 0 | — |
case-04 | pass→pass | 14,128 | 13,887 | -2% | 1 | 1 | 0% | 2,346 | 3,199 | +36% | 0 | 0 | — |
case-05 | pass→pass | 15,893 | 12,377 | -22% | 1 | 1 | 0% | 2,195 | 3,241 | +48% | 0 | 0 | — |
case-06 | pass→pass | 12,459 | 11,821 | -5% | 1 | 1 | 0% | 1,987 | 3,115 | +57% | 0 | 0 | — |
case-07 | pass→pass | 10,687 | 15,074 | +41% | 1 | 1 | 0% | 1,935 | 2,552 | +32% | 0 | 0 | — |
case-08 | pass→pass | 13,572 | 14,145 | +4% | 1 | 1 | 0% | 2,110 | 3,493 | +66% | 0 | 0 | — |
case-09 | pass→pass | 15,443 | 13,435 | -13% | 1 | 1 | 0% | 2,491 | 3,443 | +38% | 0 | 0 | — |
case-10 | pass→pass | 14,227 | 13,409 | -6% | 1 | 1 | 0% | 1,920 | 3,378 | +76% | 0 | 0 | — |
case-11 | pass→pass | 10,346 | 7,604 | -27% | 1 | 1 | 0% | 1,644 | 2,755 | +68% | 0 | 0 | — |
case-12 | fail→pass | 17,877 | 15,747 | -12% | 1 | 1 | 0% | 3,272 | 3,885 | +19% | 0 | 0 | — |
case-13 | pass→pass | 10,236 | 8,417 | -18% | 1 | 1 | 0% | 1,669 | 2,565 | +54% | 0 | 0 | — |
case-14 | pass→pass | 15,342 | 13,154 | -14% | 1 | 1 | 0% | 2,257 | 3,357 | +49% | 0 | 0 | — |
case-15 | pass→pass | 11,128 | 10,791 | -3% | 1 | 1 | 0% | 2,016 | 3,069 | +52% | 0 | 0 | — |
case-16 | pass→pass | 14,706 | 13,972 | -5% | 1 | 1 | 0% | 2,212 | 3,343 | +51% | 0 | 0 | — |
case-17 | pass→pass | 9,161 | 5,022 | -45% | 1 | 1 | 0% | 1,542 | 2,011 | +30% | 0 | 0 | — |
case-18 | pass→pass | 12,543 | 8,731 | -30% | 1 | 1 | 0% | 2,186 | 2,662 | +22% | 0 | 0 | — |
case-19 | pass→pass | 16,679 | 14,810 | -11% | 1 | 1 | 0% | 2,810 | 3,842 | +37% | 0 | 0 | — |
case-20 | pass→pass | 20,921 | 19,476 | -7% | 1 | 1 | 0% | 3,189 | 4,364 | +37% | 0 | 0 | — |
case-21 | pass→pass | 16,231 | 20,156 | +24% | 1 | 1 | 0% | 2,510 | 4,420 | +76% | 0 | 0 | — |
case-22 | pass→pass | 16,081 | 17,690 | +10% | 1 | 1 | 0% | 2,854 | 4,118 | +44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.