Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when an AI/ML selection tool uses "dynamic" models or norms that update frequently (sometimes after every administration), or when deciding how often to revalidate and update norms — Concern 7 of Tippins, Oswald & McPhail (2021). Covers the real-change-vs-instability dilemma, technical-report/documentation updates, score adjustments and grandparenting, disparate-treatment risk from candidates evaluated on different variables, applicant-pool shifts affecting validity/range restriction, and re
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -20% | 0% |
ML-derived selection procedures enable analysis that wasn't feasible before, and many vendors refresh the algorithm frequently — sometimes after every test administration. Such "dynamic" procedures update validation evidence and normative data in near real time. That creates a dilemma and several operational problems.
When a dynamic model changes, the change could reflect either:
update), or
Distinguishing the two is hard — and capitalizing on instability degrades the model.
every update to the underlying selection process requires updating the technical report (validity and normative-data characteristics). Out-of-date reports are problematic for legal defensibility and for HR records maintenance over time. (See technical-validation-report.)
already in the database, or a policy on how to treat scores produced by different processes at different times. Continual adjustment creates administration problems: a candidate qualified today may be unqualified tomorrow (or vice versa). Changes can include altering predictor weights, adding/removing predictors, and changing interpretations (e.g., adjusting cutoff scores). Grandparenting policies are administratively hard to manage in high-volume programs. (See selection-decisions-and-scoring.)
candidates are evaluated on different variables depending on when they applied. If that variation relates to a protected characteristic, there's a disparate-impact specter; even without group-level disparate impact, disparate-treatment concerns can arise — and the appearance of such treatment can trigger applicant dissatisfaction and complaints.
validity, range restriction, or range enhancement on big data — so monitor key applicant-pool characteristics (demographic, educational). Future shifts in available technologies/data/algorithms may change the applicant database and indicate the ML model should be updated.
applicant pools for adverse-impact analysis or appropriate normative groups for comparison.
Employers have always revalidated and updated norms; with AI the difference is the frequency. Traditionally, revalidation was triggered when the job changed, the test was compromised, the applicant pool shifted substantially, or enough time elapsed to question validity in a legal challenge — and it was undertaken at well-spaced intervals because it was laborious. Today's computing power makes continuous updating far less laborious, which raises a genuinely open question: how often should validation be refreshed to accommodate the nature of new applicant data?
and administrative requirements?
ai-validity-evidence · ai-selection-legal-landscape · generalizing-validity-evidence · technical-validation-report · administration-documentation (review/updating, records) · selection-decisions-and-scoring (cutoffs, norms)
Source: Tippins, Oswald & McPhail (2021), Concern: "Changes to Technologically Enhanced Systems" (Dynamic models and norms; Revalidation and norms updating).
Other measured skills in the registry, with their headline benchmark lift.