Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use FIRST when evaluating, classifying, or comparing any AI-based or technologically enhanced personnel selection tool — to separate the three independent things it combines: technologies, data, and algorithms (Tippins, Oswald & McPhail, 2021). Establishes that a technology is never "universally valid," that data range from intentional to incidental, and that ML effectiveness depends more on data quality than algorithm choice. Triggers: "evaluate an AI hiring tool", "is this video-interview/game
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 14% | 0% |
The starting lens for the whole "scientific, legal, and ethical concerns" framework. Before asking "is this AI tool any good?", decompose it into three independent parts and evaluate each separately (the "modular approach," Lievens & Sackett, 2017). Conflating them is the root error behind most overclaiming.
Examples: online games, video interviews, social media, gamification, VR. Technologies are independent from the constructs being measured and should not be confused with them (Arthur & Villado, 2008; Campbell & Fiske, 1959). Most formats can measure a wide range of constructs, and most constructs can be assessed by many technologies.
Consequence (state this in any evaluation): a new technology cannot be said to be universally valid. "Our video-interview platform is validated" is a category error — validity is a property of the inferences about the constructs measured in a specific use, not of the technology. Demand evidence about which job-relevant constructs are measured, which then informs validity and fairness.
Data vary along a continuum (Oswald, 2020):
answer).
sometimes little applicant control): social-media posts, facial movements, voice characteristics in a video, mouse clicks, response times.
Less obtrusive technologies tend to collect more incidental data, in massive amounts (game data = every click/decision/scenario; video = continuous voice and facial features; "big data" pulled from resumes, emails, social media). The more incidental the data, the more the concerns about job relevance, control, consent, and fairness intensify (see ai-candidate-data-control).
enough to appear intelligent.
used — e.g., predicting supervisory performance ratings); unsupervised learning groups people/cases into clusters with no criterion (e.g., applicants like/unlike high performers).
The "learning" happens when algorithms are first exposed to a training set; the model is judged on an independent test set (a hold-out sample, k-fold folds, or newly collected data).
There are hundreds of ML algorithms, and different algorithms often make highly similar predictions with similar overall accuracy (Domingos, 2012). So in personnel selection, the effectiveness of ML prediction/clustering is more likely driven by the availability of high-quality data than by which ML algorithm is chosen. Whether the advantages of a large number of predictors offset the disadvantages of "messy" data must be determined case by case. Accurate, well-justified predictions depend on good measurement processes and good data — not on algorithm sophistication. Don't let a vendor's algorithm story distract from data-quality and construct questions.
intentional–incidental continuum, which algorithm type — supervised/unsupervised).
measurement job-relevant, reliable, valid, and fair?
control/consent/fairness concerns.
construct-relevance story.
ai-selection-legal-landscape) and the 11 concerns.ai-selection-legal-landscape · all 11 concern skills · ai-validity-evidence · ai-input-data-and-design-audit (the audit counterpart) · validation-planning
Source: Tippins, Oswald & McPhail (2021), "New Forms of Assessment."
Other measured skills in the registry, with their headline benchmark lift.