Skip to content
Answer Stack
Open menu

Are pre-employment tests and skills assessments actually worth it, and do they predict job performance?

✓ Verified Last reviewed by AnswerStack Next review due Oct 21, 2026

Every claim is sourced below

Pre-employment tests do predict job performance, though less strongly than the field believed for most of the past 25 years. A 2022 reanalysis in the Journal of Applied Psychology showed that the range restriction corrections used across decades of meta-analyses had systematically inflated validity estimates, pulling most published figures down by roughly .10 to .20 points [1][2]. On the revised numbers, structured interviews (.42), job knowledge tests (.40), empirically keyed biodata (.38), and work sample tests (.33) rank at or above general cognitive ability (.31), which had held the top spot since 1998 [2]. Whether a given test is worth buying turns on hiring volume, whether the job has requirements you can measure before hire, what the assessment replaces, and how much subgroup difference it adds, since a procedure with adverse impact is treated as discriminatory unless it has been validated [3][5].

Do pre-employment tests actually predict job performance?

Yes, though by less than the figures still quoted in most textbooks and vendor decks. The reference point for three decades was Schmidt and Hunter's 1998 summary, which placed general mental ability at .51 and work sample tests at .54 [2]. In 2022, Paul Sackett, Charlene Zhang, Christopher Berry, and Filip Lievens found a methodological problem behind those coefficients: meta-analyses had applied range restriction corrections derived from predictive studies to concurrent studies, where restriction is usually small, and roughly 80% of these studies are concurrent [1]. Correcting heavily where little restriction exists produces overcorrection, so the published validities came out too high [1].

What the revised numbers look like

Mean operational validity fell by about .10 to .20 points across most predictors, and further for several: .21 lower for work sample tests, .20 lower for cognitive ability, .19 lower for unstructured interviews [2]. The rank order shifted as well. Structured interviews came out on top at .42, then job knowledge tests at .40, empirically keyed biodata at .38, and work sample tests at .33, with general mental ability at .31 [2]. A later meta-analysis limited to 21st century studies put cognitive ability lower still at .23, dropping it from fifth to twelfth [2].

Averages hide a large amount of variation

Mean validity is a poor guide on its own because the spread across studies is wide. Structured interviews average .42 with a standard deviation of .19, giving an 80% credibility interval of .18 to .66 [2]. A carefully built procedure lands near the top of that band and a rushed one near the bottom, so execution quality matters about as much as method choice [2].

The table compares common selection methods on the four factors that decide whether one earns a place in a hiring process. Validity figures and Black-White standardized mean differences are the revised estimates published by Sackett and colleagues [2].

Method Validity Black-White d Candidate friction Best fit
Structured interview .42, SD .19 .23 Low for candidates, high for employer Roles where trained interviewers and a fixed question set are realistic
Job knowledge test .40, SD .13 .54 Low Jobs needing the knowledge at hire, such as licensed trades
Empirically keyed biodata .38, SD .09 .33 Low Volume screening with enough hires to key against outcomes
Work sample test .33, SD .09 .67 High Experienced hiring where the work can be simulated
Assessment center .33 after revision .52 Very high Management selection
Cognitive ability test .31, or .23 recently .79 Low Trainee roles, at small weight beside other measures
Integrity test .31, SD .20 .10 Low Volume hiring where counterproductive behavior is the risk
Situational judgment test .26 .34 to .39 Moderate Entry-level roles requiring no experience
Conscientiousness questionnaire .21, or .25 contextualized -.07 Low, but distortable [9] A supporting measure inside a composite
Unstructured interview .19 .32 High for both sides Mutual-fit conversation, not scoring

Two cautions apply. Some standardized mean differences in the published table come from nonapplicant or mixed samples rather than live applicant pools, so treat them as directional, and every validity figure is for a single predictor used alone, since a composite outperforms any row shown [2].

Which methods hold up, and what is each one good for?

Structured interviews lead the ranking because of standardization rather than anything special about interviewing. Every candidate answers the same predetermined questions and every answer is scored against defined anchors, which removes the noise that leaves an unstructured conversation predicting at only .19 [2]. The cost sits with the employer, who writes job-relevant questions, builds scoring guides, and trains the panel.

Job knowledge tests

Job knowledge tests measure what a candidate already knows and rank second at .40, with a tight standard deviation of .13 [2]. They are cheap to run and straightforward to defend on content validity grounds when items map to a documented job analysis [6]. They apply only where the knowledge is required at hire rather than trained afterward, and they carry a sizable subgroup difference of .54 [2].

Empirically keyed biodata

Empirically keyed biodata scores background responses against items statistically linked to later job outcomes, and it reaches .38 [2]. Its practical edge is the low end of its credibility interval, .26 against .18 for structured interviews, so a risk-averse employer has less downside [2]. Keying an instrument empirically needs enough hires with outcome data, which is why it suits volume programs.

Work samples and assessment centers

Work samples ask a candidate to perform a slice of the job and land at .33, well below the .54 previously reported [2]. Assessment centers, effectively multi-exercise work samples for managerial roles, land at the same .33 after correction, and both consume heavy candidate and assessor time while carrying large subgroup differences of .67 and .52 [2].

Cognitive ability tests

General mental ability tests fell furthest, from .51 to .31, and to .23 in 21st century samples [2]. They stay cheap and fast and work with candidates who have no relevant experience, but they carry the largest Black-White standardized mean difference in the set, .79 [2].

Integrity tests

Integrity tests predict at .31, revised down from .41 [2]. They keep appearing in optimized composites because they pair usable validity with a subgroup difference of only .10, adding predictive power without adding adverse impact [2][7].

Situational judgment tests

Situational judgment tests predict at .26 whether scored on knowledge or on behavioral tendency, and they work with candidates who have no job experience [2]. In multi-predictor models they often receive little weight once structured interviews and biodata are in the mix [7].

Personality questionnaires

Conscientiousness measures predict at .21 overall and .25 when items are worded to a work context, and they carry essentially no subgroup difference, so they help a composite on the diversity side [2]. Faking is the standing concern, since research on distortion in high-stakes settings shows candidates with higher cognitive ability are better at shifting responses toward an ideal profile [9].

When is an assessment worth the cost, and when is it not?

Assessments earn their cost where hiring volume is high enough for a small gain in hire quality to compound, where the job has requirements you can measure before hire, and where the realistic alternative is an unstructured interview predicting at .19 [2].

Incremental validity is the number to ask for

What a test adds beyond your existing process matters more than its standalone coefficient. Across all composites of one to five predictors, mean validity comes to .51 using the pre-2022 figures and .47 using the revised ones, so multi-measure processes lost far less than any single measure did [2]. Removing cognitive ability entirely from a six-predictor composite cost .20 of validity under the old estimates and only .05 under the revised ones [2][7].

Where assessments tend not to pay

Very low volume is the first case, because a job analysis, a validation study, and ongoing monitoring carry a fixed cost that a handful of hires per year cannot absorb. The second is a mismatch between what the test measures and what the job requires, which fails on prediction and on the job-relatedness standard the EEOC applies [3]. The third is friction that removes strong candidates before anyone evaluates them, invisible in any vendor validity report. Where any of those holds, the same money spent building a real structured interview generally returns more, since it is the highest-validity method available and needs no vendor [2].

How serious is the adverse impact trade-off?

Adverse impact decides more assessment programs than validity does, because a procedure that disproportionately screens out a protected group is treated as discriminatory unless it has been validated [5]. Federal enforcement agencies generally treat a selection rate below four-fifths of the highest group's rate as evidence of adverse impact [4].

The evidence side of the trade-off

Subgroup differences vary enormously across methods and do not track validity. Cognitive ability tests carry the largest Black-White standardized mean difference at .79, then work samples at .67, job knowledge tests at .54, and assessment centers at .52 [2]. At the other end sit integrity tests at .10, structured interviews at .23, and conscientiousness at -.07 [2]. The older assumption was that reducing the weight on cognitive ability cost a great deal of validity, and the revised figures ease that dilemma. In a Pareto-optimized composite of six methods, the validity-maximizing solution reaches .61 at an adverse impact ratio of .42, and zeroing out cognitive ability drops validity only to .56 while raising the ratio to .67, so excluding it barely moves validity yet substantially reduces adverse impact [7][8].

What the trade-off still costs

Pushing that same model past the .80 ratio threshold requires validity to fall further, to .48, so the tension has narrowed rather than disappeared [7]. Where two procedures are substantially equally valid, the guidelines direct employers to the one with lesser adverse impact [5], and the EEOC lists searching for a less discriminatory alternative among its best practices [3]. An employer relying on a test must show it is job-related and consistent with business necessity, supported by a criterion-related, content, or construct validity study, and a vendor technical manual does not transfer that responsibility [3][6]. This is general information about federal standards, not legal advice, and state and local rules add obligations these sources do not cover.

What do assessments cost in money and in candidates?

Published assessment pricing is generally a subscription plus per-candidate credits rather than a flat per-test fee. TestGorilla's own pricing page lists a free tier with 10 credits per month, a Core plan at $142 per month billed annually, a Plus plan from $400 per month billed annually, and custom enterprise terms, with a skills test consuming one credit per candidate [11]. Those are vendor-published figures as of July 2026 and are not independently audited.

Candidate drop-off is the hidden cost line

The cost that rarely reaches a business case is the candidate you never evaluate. Criteria Corp reports, from an analysis of roughly 500,000 of its own tests, that employers can expect a completion rate near 80% for assessments up to 40 minutes, falling toward 60% once total testing runs beyond an hour [10]. That is a vendor claim about a vendor's own platform, and no independent benchmark of comparable scale was available to corroborate it. A 40% loss at the assessment stage outweighs the difference between most competing instruments, so calculate cost per completed assessment rather than cost per license, then weigh it against the validity the instrument adds on top of a structured interview [2].

What should you check before buying any assessment?

Ask for the technical report before you ask for a demo. A credible vendor supplies validation evidence, subgroup statistics, and the samples behind them, and one that cannot is asking you to carry a legal obligation on trust [3].

The checks worth running, in order:

  • Start from a documented job analysis, since the EEOC's standard is whether the procedure is job-related for the position and appropriate to your purpose [3].
  • Ask which validation strategy supports the instrument; the Uniform Guidelines accept criterion-related, content, or construct validity studies [6].
  • Ask whether the underlying samples were concurrent or predictive and how range restriction was handled, the exact issue that inflated older estimates [1].
  • Ask for the credibility interval, not the mean validity alone, because .42 ranging from .18 to .66 is a different purchase than .40 with a narrow band [2].
  • Request subgroup data, then calculate your own selection rates by group against the four-fifths standard once the tool is live [4].
  • Compare at least one lower-impact alternative of comparable validity, since the guidelines point you toward the equally valid procedure with less adverse impact [5].
  • Measure completion time and completion rate during a pilot, then decide whether the drop-off is acceptable at your volume [10].
  • Confirm the accommodation path for candidates with disabilities before launch, because Americans with Disabilities Act obligations apply regardless of who built the tool [3].

What pre-employment testing is not

It is not a lie detector

Integrity tests are written questionnaires and sit outside polygraph law, though the two get conflated. The Employee Polygraph Protection Act makes it unlawful for most private employers to "directly or indirectly, to require, request, suggest, or cause any employee or prospective employee to take or submit to any lie detector test" [12].

It is not a medical examination

Medical inquiries and physical examinations fall under separate rules with timing restrictions, distinct from cognitive, personality, and skills testing [3]. A personality inventory that strays into questions likely to reveal a mental health condition can cross that line.

It is not a personality type sorter

Instruments that assign candidates to types are built for development conversations, and none of the validity evidence here applies to them. The measured predictor in the research is trait conscientiousness, not a type label [2].

It is not a substitute for a job analysis

Buying a validated instrument does not make it valid for your job, since validation is position-specific and the employer holds that burden [3][6].

This answer draws on the primary industrial and organizational psychology literature rather than assessment marketing, because the validity figures most often quoted in vendor material trace back to a 1998 summary that a 2022 reanalysis has since revised downward [1][2]. Where a number comes from a vendor it is labeled as a vendor claim and the vendor's own page is cited, since no independent audit of completion rates or per-candidate pricing at comparable scale was available. Legal material comes from the EEOC and from the Uniform Guidelines text hosted by Cornell's Legal Information Institute, and it is general information rather than legal advice. HR practitioners, industrial and organizational psychologists, and vendors holding local validation studies, subgroup statistics, or completion data that contradict anything here are invited to send it so the record can be corrected.

This answer was written and reviewed by the AnswerStack Editorial Team, which has no commercial stake in the products, companies, or methods discussed. Every claim is cited inline and verified on the dates shown.

Sources

Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range

Journal of Applied Psychology (Sackett, Zhang, Berry & Lievens, 2022) via PubMed

Primary source Verified Jul 21, 2026 Supports: Range restriction artifact distributions produced substantial overcorrection; validity of many selection procedures substantially overestimated; structured interviews emerged as top-ranked; mean estimates reduced by .10 to .20 points

“Key findings are that most of the same selection procedures that ranked high in prior summaries remain high in rank, but with mean validity estimates reduced by .10-.20 points. Structured interviews emerged as the top-ranked selection procedure.”

Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors

Industrial and Organizational Psychology 16(3), Cambridge University Press (Sackett, Zhang, Berry and Lievens, 2023)

Primary source Verified Jul 21, 2026 Supports: Table 1 revised operational validity estimates, standard deviations and Black-White d values for every method cited; Schmidt and Hunter (1998) comparison figures; roughly 80% of studies concurrent; drops of .21, .20 and .19 for work samples, cognitive ability and unstructured interviews; structured

“Structured interviews, which top our list in terms of a mean validity of .42 have an 80% credibility interval ranging from .18 to .66. Thus, the validity of structured interviews should really be viewed as '.42, plus or minus .24'.”

Employment Tests and Selection Procedures

U.S. Equal Employment Opportunity Commission

Primary source Verified Jul 21, 2026 Supports: Types of test covered; Title VII disparate treatment and disparate impact; job-related and consistent with business necessity standard; ADA restrictions on medical inquiries and examinations; employer responsibility to validate even where a vendor supplies documentation; best practice of seeking an

“Employers should ensure that employment tests and other selection procedures are properly validated for the positions and purposes for which they are used.”

29 CFR 1607.4: Information on impact (Uniform Guidelines on Employee Selection Procedures)

Legal Information Institute, Cornell Law School

Primary source Verified Jul 21, 2026 Supports: The four-fifths (80 percent) rule as evidence of adverse impact, and the requirement to track selection rates by race, sex and ethnic group

“A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact.”

29 CFR 1607.3: Discrimination defined, relationship between use of selection procedures and discrimination

Legal Information Institute, Cornell Law School

Primary source Verified Jul 21, 2026 Supports: Adverse impact treated as discriminatory absent validation; duty to prefer the substantially equally valid procedure with lesser adverse impact

“Where two or more selection procedures are available which serve the user's legitimate interest in efficient and trustworthy workmanship, and which are substantially equally valid for a given purpose, the user should use the procedure which has been demonstrated to have the lesser adverse impact.”

29 CFR 1607.5: General standards for validity studies

Legal Information Institute, Cornell Law School

Primary source Verified Jul 21, 2026 Supports: Criterion-related, content and construct validity studies are the accepted validation strategies under the Uniform Guidelines

“For the purposes of satisfying these guidelines, users may rely upon criterion-related validity studies, content validity studies or construct validity studies, in accordance with the standards set forth in the technical standards of these guidelines, section 14 below.”

Insights from an updated personnel selection meta-analytic matrix: Revisiting general mental ability tests' role in the validity-diversity tradeoff (author manuscript)

Berry, Lievens, Zhang and Sackett, Journal of Applied Psychology 109(10), 1611-1634, author copy hosted by Filip Lievens

Primary source Verified Jul 21, 2026 Supports: Updated matrix validities and Black-White d values; Pareto-optimized composite reaching .61 validity at an adverse impact ratio of .42; zeroing cognitive ability yields .56 validity and a .67 ratio; reaching a ratio above .80 drops validity to .48; integrity tests and structured interviews carry the

“The validity-maximizing solution has a validity composite of .61 and an adverse impact ratio of .42... When the weight of GMA tests gets dropped to 0, validity only decreases to .56 and the adverse impact ratio increases to .67.”

Insights From an Updated Personnel Selection Meta-Analytic Matrix: Revisiting General Mental Ability Tests' Role in the Validity-Diversity Trade-Off

Experts@Minnesota, University of Minnesota

Supporting Verified Jul 21, 2026 Supports: Publication record confirming the 2024 Journal of Applied Psychology citation and the conclusion that excluding cognitive ability tests has little to no effect on validity while substantially decreasing adverse impact

“Our results lead to the conclusion that excluding GMA tests generally has little to no effect on validity, but substantially decreases adverse impact.”

The "g" in Faking: Doublethink the Validity of Personality Self-Report Measures for Applicant Selection

Frontiers in Psychology (Geiger, Olderbak, Sauter and Wilhelm, 2018), via PubMed Central

Independent Verified Jul 21, 2026 Supports: Candidates with higher cognitive ability are better able to shift personality self-report responses toward an ideal profile, which complicates interpretation of a single personality score

“Crystallized intelligence showed the strongest effect predicting faking ability.”

Best Practices for Implementing Pre-Employment Testing

Criteria Corp (vendor)

Supporting Verified Jul 21, 2026 Supports: Vendor claim, based on the vendor's analysis of roughly 500,000 of its own tests, that completion runs near 80% for assessments up to 40 minutes and falls toward 60% beyond one hour; vendor guidance on benchmarking scores against incumbent employees

“organizations can expect a completion rate of around 80% for assessments up to 40 minutes”

Pricing

TestGorilla

Primary source Verified Jul 21, 2026 Supports: Vendor-published plan structure and prices as of July 2026: free tier with 10 monthly credits, Core at $142 per month billed annually, Plus from $400 per month billed annually, custom enterprise terms, and credit consumption per candidate

“10 free credits every month”

29 U.S. Code 2002: Prohibitions on lie detector use

Legal Information Institute, Cornell Law School

Primary source Verified Jul 21, 2026 Supports: Employee Polygraph Protection Act prohibition on requiring, requesting or suggesting a lie detector test and on using its results

“directly or indirectly, to require, request, suggest, or cause any employee or prospective employee to take or submit to any lie detector test”

Revision history

2 revisions since publication
v1.1 Reviewed and re-verified.
v1.0 Published after editorial review.