Skip to content
Answer Stack
Open menu

Does AI resume screening software actually work, and is it worth it (bias, accuracy)?

✓ Verified Last reviewed by AnswerStack Next review due Oct 21, 2026

Every claim is sourced below

AI resume screening reliably reproduces the ranking it was trained to reproduce, and independent audits show it does so with measurable demographic disparities: a University of Washington audit of embedding models ranking 500 real resumes against 500 job descriptions found white-associated names preferred in 85.1 percent of cases, female-associated names in 11.1 percent, and Black male-associated names disadvantaged in up to 100 percent of comparisons [1]. Whether it is worth buying turns on what the model was trained to predict, because vendors generally let the client pick the outcome, so a tool can be accurate at copying prior screening while adding nothing to hiring quality [4]. Corrected meta-analytic estimates also cut the validity of most selection procedures by .10 to .20 and left the structured interview highest ranked [7]. In New York City the tool needs a bias audit from the past year, published impact ratios, and candidate notice 10 business days before use [8]. This is general information, not legal advice.

What does the independent research show about AI resume screening?

The strongest evidence comes from resume audit studies, and it points where human-hiring audits have long pointed: identical resumes rank differently depending on the name at the top. Kyra Wilson and Aylin Caliskan of the University of Washington ran more than 500 public resumes against 500 job descriptions across nine occupations using Massive Text Embedding models, the class behind semantic resume matching in commercial tools [1]. White-associated names were preferred in 85.1 percent of cases against 11.1 percent for female-associated names, and Black male-associated names were disadvantaged in up to 100 percent of comparisons [1].

A follow-up experiment tested what happens with a person in the loop. Across 528 participants and 1,526 screening decisions, people shown no recommendations picked across racial groups at roughly equal rates, but when the model favored a group they picked from it up to 90 percent of the time, even when they called the recommendations poor quality [2].

Newer models complicate the picture rather than closing it. A June 2026 paired-resume audit of fourteen language models, at 24,024 paired postings each, found the one 2023-vintage model reproduced a pro-white callback gap of 2.12 percentage points, while every model from 2024 onward showed no significant gap or a reversal of up to 3.01 points [3].

The case people cite most is older and different in kind. Amazon started building a system in 2014 to score applicants from one to five stars, trained on ten years of resumes it had received; because most came from men, it learned to penalize the word "women's" and the names of certain all-women colleges, and Amazon scrapped it in 2017 [6]. That came from supervised learning on hiring records, not from a language model.

Four technologies get sold under one label, and the bias and accuracy story differs for each. Most coverage treats them as one thing, which is why it rarely helps.

Technology How it ranks or filters Where bias enters What accuracy means here
Keyword and boolean matching Literal term presence, plus knockout rules on fields like years of experience Requirements written into the query, and applicant vocabulary differences Whether the query matched the text, not anything about the candidate
Statistical ranking trained on past outcomes Scores applicants against a target the employer chose, learned from past records The past decisions, plus proxies that survive removal of protected attributes [5] Agreement with the target, usually a prior decision, not performance [4]
Language model or embedding evaluation of free text Semantic similarity to a job description, or a model's written judgment Associations carried in the pretrained model, including name-based ones [1][3] Agreement with a rater or a similarity score, not a validated prediction
Structured parsing and field extraction Reads a resume into database fields with no score attached Extraction failures on unusual formats, not a scoring rule Whether the fields came out right, which you can check by hand [8]

How does each screening method actually behave?

Keyword and boolean matching

Keyword filters do what the query says and nothing more, so their errors are a recruiter's errors made faster: a string requiring "CPA" drops the candidate who wrote "Certified Public Accountant." Bias here is written into the requirements rather than learned from data, so it is the easiest kind to fix. Pull the live query on an open requisition and count how many applicants each rule removes.

Statistical ranking trained on past hiring outcomes

This category produced the Amazon failure, and it is where bias is structural rather than incidental. The model learns from records of who was previously advanced or rated well, so patterns in those decisions become patterns in the score, and stripping name, gender, and race from the inputs does not remove them because proxies survive in schools, employers, and employment gaps [5]. Upturn documented a screener that had learned the name "Jared" and high school lacrosse predicted success, a real statistical association with no causal link to the work [5].

Language model and embedding evaluation of free text

These tools carry associations from pretraining rather than from your hiring history, which is why disparities appear with no employer data involved: the Washington audit ran on off-the-shelf embedding models with no fine-tuning and the name gaps appeared anyway [1]. Those associations move between releases, so treat the model version and prompt as a fixed configuration and re-audit whenever either changes [3].

Structured parsing and field extraction

Parsing carries the lowest risk of the four because it produces no score and expresses no preference. New York City's rule draws the same line, excluding from its definition of a simplified output tools that only translate or transcribe existing text, such as converting a resume from a PDF [8]. The real risk is quiet data loss, since a mangled skills section becomes a thin record every downstream filter inherits.

Does AI resume screening actually work?

"Works" has no agreed meaning in this market, and that is the first thing to settle before reading any accuracy claim. A screening model is trained against a target variable someone chose, and in commercial tools that choice usually belongs to the buyer. Raghavan, Barocas, Kleinberg, and Levy reviewed 18 vendors of algorithmic pre-employment assessments and found vendors leaving it to clients to decide what outcomes to predict, naming performance reviews, sales numbers, and retention time as examples [4]. Vendor websites generally do not make clear whether models are validated or how validation data was selected [4].

So a tool can post an excellent accuracy figure and still be worth nothing to you. If the label it learned was "this applicant was advanced to a phone screen last year," high accuracy means it agrees with last year's recruiters. It has automated the existing screen, including whatever that screen was getting wrong, with no evidence about who does the job well.

A tool trained on real performance data still meets a ceiling that predates this software. Sackett and colleagues re-examined the meta-analytic validity literature and concluded that range restriction corrections had substantially overstated the validity of most selection procedures, cutting mean estimates by .10 to .20 points, with the structured interview highest ranked [7]. Nothing in personnel selection predicts job performance with anything close to precision, so a tool claiming to find the best candidate is claiming something the field has never shown.

How would you know whether your tool is biased?

The measurement standard already exists in federal enforcement practice: compare selection rates across groups and flag any group below four-fifths of the highest rate. Under 29 CFR 1607.4(D), such a rate "will generally be regarded by the Federal enforcement agencies as evidence of adverse impact" [12], and the Uniform Guidelines require validity evidence only where a procedure adversely affects a group's opportunities [13].

New York City turned that arithmetic into a filing requirement. A bias audit under the DCWP rule must calculate the selection rate and impact ratio for every EEO-1 category, separately for sex, for race and ethnicity, and for intersectional categories, and must report how many assessed people fell into an unknown category [8]. Where the tool scores rather than selects, scoring rate above the sample median replaces selection rate [8].

How often any of this happens is a separate question. A 2024 study at the ACM Conference on Fairness, Accountability, and Transparency sent 155 investigators to 391 employer websites and found audit reports for 18 employers and transparency notices for 13 [9]. The authors call that "null compliance" rather than non-compliance, since employers decide for themselves whether a tool is in scope [9].

What do the rules require as of July 2026?

Three jurisdictions impose specific obligations on employers using automated tools to screen candidates, while the federal baseline stays ordinary discrimination law rather than an AI statute. This is general information, not legal advice; confirm current requirements with counsel before deploying anything.

New York City

Local Law 144, implemented by DCWP rules effective July 5, 2023, bars use of an automated employment decision tool unless it has had a bias audit within the past year, a summary of results is published on your site before use, and notice reached candidates at least 10 business days beforehand with instructions for requesting an alternative process or accommodation [8]. The definition reaches any tool whose simplified output is relied on alone, weighted more heavily than any other criterion, or used to overrule human conclusions, so a ranking tool sitting beside recruiter judgment can still be covered [8].

Illinois

HB 3773 became Public Act 103-0804 and took effect January 1, 2026, amending the Illinois Human Rights Act to bar artificial intelligence that discriminates on the basis of a protected class or that uses zip codes as a proxy for one, and requiring notice to workers when AI is used in hiring, promotion, discipline, and other employment decisions [10].

Colorado

Colorado's position moved twice, so older summaries are unreliable. SB26-189, signed May 14, 2026, repeals the 2024 Colorado AI Act and reenacts it as a disclosure and rights framework effective January 1, 2027 [11]. Deployers will have to notify workers before covered technology influences an employment decision, give a plain-language explanation of an adverse outcome within 30 days, and allow meaningful human review [11].

Federal

The Uniform Guidelines on Employee Selection Procedures still govern, and they reach any procedure used to make an employment decision, including application screening [13]. The EEOC's 2023 technical assistance document on adverse impact in algorithmic selection no longer resolves on eeoc.gov, so anchor your analysis to the Uniform Guidelines and 29 CFR 1607 instead [12][13].

When is AI resume screening defensible, and when is it not?

AI screening holds up where applicant volume exceeds human capacity, requirements are concrete, the tool ranks instead of rejecting, and someone recalculates impact ratios on a set cadence.

Conditions where it holds up

High-volume hourly and entry-level requisitions are the clearest case, because the realistic alternative is a recruiter skimming a thousand resumes at a few seconds each. A ranking that surfaces candidates for a human to read, with the full pool still reachable, changes the order of review instead of the outcome. A forklift certification is a fact you can verify, and a tool sorting on facts is auditable in a way a tool scoring "culture fit" is not [4]. The configuration that survives scrutiny is ranking, plus human review of the list, plus a standing adverse impact calculation [12].

Conditions where it is hard to defend

Small requisitions rarely justify the exposure, since forty applications are readable by a person and the audit overhead outweighs the time saved. Subjective and senior roles fit poorly, because the qualities that matter are the ones a model reaches through proxies, where the documented disparities live [1][5]. Automatic rejection without review turns a ranking error into a final decision, and it sits inside the NYC definition of a tool relied on alone [8]. Without category-level rates on your applicant flow, you cannot see whether the four-fifths threshold is being crossed [8][12].

What should you require from a vendor before you buy?

Ask for five things in writing, and treat a refusal as its own answer.

The most recent bias audit, with impact ratios by category

Request the audit output itself rather than a compliance badge: selection or scoring rates and impact ratios for each sex, race and ethnicity, and intersectional category, plus the audit date, data source, and count excluded as unknown [8]. An audit on the vendor's pooled data across many employers is not the same as one on your applicant flow, and the NYC rule expects the auditor to get your own historical data [8].

The target variable and the training population

Ask what the model was trained to predict and over which people, in one sentence each. Vendors hand the outcome choice to the client and rarely publish validation methodology, so the answer will not be on the website [4]. If the target is a past decision or rating, the tool measures agreement with prior screening, and should be judged on that.

An explanation you could give a rejected candidate

Ask the vendor to produce, for one real applicant, the reason that applicant ranked where they did. Colorado will require a plain-language explanation of an adverse decision within 30 days plus meaningful human review [11], so a tool that cannot generate one leaves that obligation to you.

The model version and the change policy

Language-model behavior varies by release, and the fourteen-model comparison found demographic gaps flipping direction across generations [3]. Get the version identifier, the change-notice policy, and a written commitment to re-audit after an upgrade.

The evidence behind any performance number

Performance figures are marketing until the underlying study is on the table. Workday markets HiredScore AI for Recruiting with a claimed "54% increase in recruiter capacity within 10 months of launch" and "unbiased, AI-driven candidate grading" [14]; both are vendor claims, not published validation of screening accuracy. Ask which customers, over what period, against what baseline.

This record was assembled from primary research papers and primary legal text rather than vendor marketing, because the vendor literature on this question is close to uniformly promotional. The audit studies were opened and read directly, and the statutes and rules were read in the original rather than through law firm summaries, which is why the Colorado section reflects the May 2026 replacement law instead of the 2024 act much secondary coverage still describes. One document others cite often, the EEOC's 2023 technical assistance on algorithmic selection, no longer resolves and is deliberately absent below. HR practitioners running these tools, and vendors who believe a claim here misstates their product, are invited to send corrections with evidence. Bias audit results from real applicant flows would be especially useful, since almost none are public.

This answer was written and reviewed by the AnswerStack Editorial Team, which has no commercial stake in the products, companies, or methods discussed. Every claim is cited inline and verified on the dates shown.

Sources

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

Kyra Wilson and Aylin Caliskan, University of Washington (AAAI/ACM AIES 2024)

Primary source Verified Jul 21, 2026 Supports: 85.1 percent preference for white-associated names, 11.1 percent for female-associated names, up to 100 percent disadvantage for Black male-associated names; 500+ resumes, 500 job descriptions, nine occupations, Massive Text Embedding models with no fine-tuning

“Using that framework, we then perform a resume audit study to determine whether a selection of Massive Text Embedding (MTE) models are biased in resume screening scenarios.”

No Thoughts Just AI: Biased LLM Hiring Recommendations Alter Human Decision Making and Limit Human Autonomy

Kyra Wilson, Mattea Sim, Anna-Maria Gueorguieva, Aylin Caliskan

Independent Verified Jul 21, 2026 Supports: 528 participants and 1,526 resume screening decisions; equal selection without AI recommendations; selection of the favored group up to 90 percent of the time with biased recommendations; effect persisting among participants who rated the recommendations poorly

“When interacting with AI favoring a particular group, people select those candidates up to 90% of the time.”

Can LLMs Hire Fairly? Racial Bias in Resume Screening

Zhenyu Gao, Wenxi Jiang, Yutong Yan (arXiv preprint, June 2026)

Independent Verified Jul 21, 2026 Supports: Fourteen models, 24,024 paired postings per model; 2023-vintage model shows a +2.12 percentage point pro-white gap; 2024 and later models show null gaps or reversals up to 3.01 points; same generational pattern on gender

“The sole 2023-vintage model reproduces the pro-White callback gap.”

Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices

Raghavan, Barocas, Kleinberg and Levy (ACM FAccT 2020)

Independent Verified Jul 21, 2026 Supports: 18 vendors of algorithmic pre-employment assessments reviewed; eight build assessments from client employee data; vendors leave target variable choice to clients; websites generally do not make validation methodology clear

“Vendors in general leave it up to clients to determine what outcomes they want to predict, including, for example, performance reviews, sales numbers, and retention time.”

Help Wanted: An Examination of Hiring Algorithms, Equity, and Bias

Upturn

Independent Verified Jul 21, 2026 Supports: Models trained on historical hiring records reproduce those patterns; removing protected attributes does not prevent proxy effects; the Jared and high school lacrosse example of a meaningless learned correlation

“Removing or obscuring sensitive factors like gender and race will not prevent predictive models from reflecting patterns of bias.”

Amazon ditched AI recruitment software because it was biased against women

MIT Technology Review

Independent Verified Jul 21, 2026 Supports: Amazon system started 2014, scrapped 2017, trained on ten years of submitted resumes, penalized resumes containing the word women's and the names of certain all-women colleges

“The company lost confidence that the program was indeed gender neutral in all other areas.”

Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range

Journal of Applied Psychology (Sackett, Zhang, Berry & Lievens, 2022) via PubMed

Independent Verified Jul 21, 2026 Supports: Prior validity estimates for selection procedures substantially overstated; revised estimates reduced by .10 to .20; structured interviews ranked highest

“Key findings are that most of the same selection procedures that ranked high in prior summaries remain high in rank, but with mean validity estimates reduced by .10-.20 points. Structured interviews emerged as the top-ranked selection procedure.”

Notice of Adoption of Final Rule: Automated Employment Decision Tools (6 RCNY Subchapter T)

New York City Department of Consumer and Worker Protection

Primary source Verified Jul 21, 2026 Supports: AEDT definition and the substantially assist test; bias audit contents including selection rate, impact ratio, sex, race/ethnicity and intersectional categories, unknown-category count, scoring rate above median, the 2 percent exclusion; published results and six-month posting; 10 business days noti

“Provide notice on the employment section of its website in a clear and conspicuous manner at least 10 business days before use of an AEDT.”

Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability

Wright, Muenster, Vecchione, Qu, Cai, Smith, Metcalf and Matias (ACM FAccT 2024)

Independent Verified Jul 21, 2026 Supports: 155 investigators searched 391 employer websites; 18 posted audit reports and 13 posted transparency notices; definition and argument for null compliance

“Null compliance describes a state in which the absence of evidence of compliance cannot be ascertained as non-compliance.”

Illinois HB 3773 bill status, Public Act 103-0804

Illinois General Assembly

Primary source Verified Jul 21, 2026 Supports: Public Act 103-0804, effective January 1, 2026; amends the Illinois Human Rights Act on AI discrimination, zip codes as a proxy, and employee notice

“Bill status page confirming Public Act 103-0804 and its January 1, 2026 effective date; the act bars AI use that discriminates on a protected class or uses zip codes as a proxy for one.”

SB26-189 Automated Decision-Making Technology

Colorado General Assembly

Primary source Verified Jul 21, 2026 Supports: Signed May 14, 2026; repeals and reenacts the 2024 Colorado AI Act; effective January 1, 2027; deployer notice, 30-day plain-language explanation of adverse decisions, right to meaningful human review, attorney general enforcement with a 60-day cure period

“Repeals and reenacts those provisions with new requirements.”

29 CFR 1607.4: Information on impact (Uniform Guidelines on Employee Selection Procedures)

Legal Information Institute, Cornell Law School

Primary source Verified Jul 21, 2026 Supports: Four-fifths rule as evidence of adverse impact; requirement to keep records of impact by race, sex, and ethnic group

“A selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact.”

Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures

U.S. Equal Employment Opportunity Commission

Primary source Verified Jul 21, 2026 Supports: Scope of employee selection procedures including application form screening; validity evidence required only where a procedure adversely affects a group

“The Uniform Guidelines require users to produce evidence of validity only when the selection procedure adversely affects the opportunities of a race, sex, or ethnic group for hire, transfer, promotion, retention or other employment decision.”

HiredScore AI for Recruiting

Workday

Supporting Verified Jul 21, 2026 Supports: Example of a live vendor claim: 54 percent increase in recruiter capacity within 10 months of launch, and unbiased AI-driven candidate grading. Vendor claim, not independent validation.

“54% increase in recruiter capacity within 10 months of launch.”

Revision history

2 revisions since publication
v1.1 Reviewed and re-verified.
v1.0 Published after editorial review.