Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when designing, conducting, or evaluating a criterion-related validity study — demonstrating an empirical relationship between selection-procedure (predictor) scores and work-relevant criteria. Covers predictive vs. concurrent designs, criterion development (relevance/contamination/deficiency/ reliability/bias), predictor choice, participant sampling, statistical power, data analysis, corrections for range restriction and unreliability, and combining predictors/criteria. Triggers: "criterion
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-19 | ✓→✓ | = Same ✓ | 110% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 52% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 53% | 0% |
Evidence that scores on a selection procedure are statistically related to one or more measures of work-relevant behavior or outcomes. The strongest form of "does it predict?" evidence — but also the most demanding on sample size, criterion quality, and analysis.
Three things determine whether a criterion-related study is even sensible:
most important property.
from the expected effect size, the statistic, and the chosen alpha. Range restriction and criterion unreliability inflate the corrected coefficient but also inflate its standard error — so distinguish observed from corrected coefficients when judging power. Underpowered studies are a chronic failure mode; give Type II error equal attention to Type I.
If any of these is badly deficient, a criterion-related strategy may not be feasible — consider content-based-validation or generalizing-validity-evidence.
retention interval or once performance stabilizes).
For stable cognitive abilities, predictive and concurrent estimates tend to be comparable. For noncognitive self-reports (personality, interests, SJTs) and experience-based measures (biodata), the designs can diverge — e.g., applicant faking motivation differs from incumbents'; biodata responses may reflect on-the-job experience. So results don't automatically generalize across designs or predictor types. Match the design's inference to your use.
Other design choices that matter: the basis on which sample members were selected, the population sampled (applicants vs. recent hires vs. fully experienced), and whether you're predicting higher-level work (acceptable if a substantial share advance and you use criteria at both the hire level and the higher level).
Choose criteria for relevance, freedom from contamination, and reliability — not availability or convenience. Criteria should represent important organizational, team, or individual outcomes.
success, turnover, OCBs, advancement). Need not be all-inclusive, but the link to the proposed use must be clear and rationale documented.
territory, rater knowledge of predictor scores, shift, location, rater attitudes). Minimize via standardized administration; measure and statistically control contaminants where possible.
to performance). If the criterion can't cover the full domain, state what is omitted and the implication for the inference.
subgroups. Cannot be detected from criterion scores alone; anticipate and guard against it with professional judgment.
generalize across and design to estimate the matching reliability. Internal-consistency estimates may be inadequate for ratings (they ignore rater and time variance).
Common criterion types: supervisory performance ratings (most common; ratings collected for research are preferable to administrative ratings — Jawahar & Williams, 1997), other performance indices, archival/HRIS data (verify accuracy, alignment, consistency, and data-privacy compliance before use).
the study. Verify serendipitous findings (especially from small samples) by independent replication.
familiarity or bias.
construct) to avoid uninterpretable comparisons.
faking — unstructured interviews and unproctored internet tests are higher risk) and predictor reliability (estimate it for the conditions of intended use).
document the conceptual/methodological basis, provide cross-validation evidence, and ensure the algorithm does not introduce systematic bias against subgroups.
motivation, ability, experience). Convenience samples are discouraged to the extent they're unrepresentative.
credible evidence of potential bias and sufficient data (adequate power and precision) for the proper analysis — see fairness-and-bias-analysis. A subgroup too small to analyze cannot be compared until more data exist.
Don't pick a method because the software is handy; if you delegate analysis, you retain responsibility for its suitability and accuracy.
predictor–criterion relationships; describe distributions (central tendency, variance) and interrelationships.
population value because of range restriction and criterion unreliability; apply suitable bivariate/multivariate corrections when an appropriate estimate is available.
unreliable predictor).
time variance.
coefficients don't apply. Report both corrected and uncorrected values, and use procedures built for corrected coefficients when testing significance / forming CIs. State explicitly when a coefficient is theoretical and not the operational validity.
rational weights, etc. Effective weights ≠ nominal weights (they depend on variances/covariances), especially when predictors are differentially range-restricted.
expected mean criterion standing, and subgroup differences — see selection-decisions-and-scoring.
against capitalization on chance, especially in small samples. Unit/rational weights don't shrink.
Interpret against the cumulative research literature. Unusual findings (suppressors, moderators, nonlinearity, configural scoring, differential weighting of highly correlated predictors) are suspect — require a very large sample or replication before acting on them.
work-analysis · validation-planning · generalizing-validity-evidence · fairness-and-bias-analysis · selection-decisions-and-scoring · technical-validation-report
Source: Principles (5th ed., 2018), "Sources of Validity Evidence → Criterion-Related Evidence" and "Operational Considerations → Selecting Criterion Measures / Data Analyses."
Other measured skills in the registry, with their headline benchmark lift.