Structured interviews lead the ranking because of standardization rather than anything special about interviewing. Every candidate answers the same predetermined questions and every answer is scored against defined anchors, which removes the noise that leaves an unstructured conversation predicting at only .19 [2]. The cost sits with the employer, who writes job-relevant questions, builds scoring guides, and trains the panel.
Job knowledge tests
Job knowledge tests measure what a candidate already knows and rank second at .40, with a tight standard deviation of .13 [2]. They are cheap to run and straightforward to defend on content validity grounds when items map to a documented job analysis [6]. They apply only where the knowledge is required at hire rather than trained afterward, and they carry a sizable subgroup difference of .54 [2].
Empirically keyed biodata
Empirically keyed biodata scores background responses against items statistically linked to later job outcomes, and it reaches .38 [2]. Its practical edge is the low end of its credibility interval, .26 against .18 for structured interviews, so a risk-averse employer has less downside [2]. Keying an instrument empirically needs enough hires with outcome data, which is why it suits volume programs.
Work samples and assessment centers
Work samples ask a candidate to perform a slice of the job and land at .33, well below the .54 previously reported [2]. Assessment centers, effectively multi-exercise work samples for managerial roles, land at the same .33 after correction, and both consume heavy candidate and assessor time while carrying large subgroup differences of .67 and .52 [2].
Cognitive ability tests
General mental ability tests fell furthest, from .51 to .31, and to .23 in 21st century samples [2]. They stay cheap and fast and work with candidates who have no relevant experience, but they carry the largest Black-White standardized mean difference in the set, .79 [2].
Integrity tests
Integrity tests predict at .31, revised down from .41 [2]. They keep appearing in optimized composites because they pair usable validity with a subgroup difference of only .10, adding predictive power without adding adverse impact [2][7].
Situational judgment tests
Situational judgment tests predict at .26 whether scored on knowledge or on behavioral tendency, and they work with candidates who have no job experience [2]. In multi-predictor models they often receive little weight once structured interviews and biodata are in the mix [7].
Personality questionnaires
Conscientiousness measures predict at .21 overall and .25 when items are worded to a work context, and they carry essentially no subgroup difference, so they help a composite on the diversity side [2]. Faking is the standing concern, since research on distortion in high-stakes settings shows candidates with higher cognitive ability are better at shifting responses toward an ideal profile [9].