AI accuracy by field · Hiring

Is AI accurate for hiring?

Not the way it is usually deployed. AI resume screeners are not fabricating facts the way a chatbot invents case law. They are quietly reproducing the biases in the data they learned from and presenting the result as a neutral score. The best-studied models favor white-associated and male-associated names, reject candidates over a certain age, and filter out qualified people whose resumes do not match a keyword template. AI can help you organize and summarize an applicant pool, but no automated ranking is safe to act on until a human has checked who it is quietly screening out.

Try the free claim checker

Confident, scalable, and biased in ways you cannot see

In the most rigorous audit of its kind, University of Washington researchers had large language models rank more than 550 real resumes against 500 real job postings, roughly three million comparisons in all. The models preferred white-associated names 85 percent of the time over Black-associated names, and male-associated names 52 percent of the time against female-associated names at 11 percent. Across every comparison, they never once preferred a Black man's name to a white man's.

This is a different failure mode from a chatbot that makes things up. A hiring model rarely invents a fact, it assigns a number. The danger is that the number looks objective and arrives at a scale no human review can match, so a biased ranking becomes thousands of quiet rejections before anyone notices. The models are also uncalibrated: a match score of 82 versus 78 implies a precision they do not have, and swapping only the name on an otherwise identical resume can move a candidate from the top of the list to the bottom.

85%

of the time, leading language models asked to rank identical resumes preferred a white-associated name over a Black-associated one. They never once preferred a Black man's name to a white man's.

Wilson and Caliskan, University of Washington (AIES 2024)

Why AI fails at hiring specifically

A resume-screening model learns from historical hiring data, which is a record of who a company hired and promoted in the past. If that history skews toward certain names, schools, genders, or ages, the model treats those patterns as signals of quality and reproduces them. Amazon's scrapped tool learned from a decade of mostly male applicants and taught itself to penalize the word women's. A general-purpose model does not even need a company's data to do this, because the association between names and demographics is already baked into what it read on the internet. The output is not a lie. It is a discriminatory pattern wearing the costume of a neutral score.

The second problem is false confidence with no ground truth. There is no verifiable correct answer to who is the best candidate, so a model cannot be checked against reality the way a fabricated citation can. It cannot reliably explain why it ranked one person over another, its rankings flip on trivial changes, and it presents all of it with the same fluent certainty whether it is fair or catastrophically biased. That is exactly why New York City's Local Law 144 now requires an independent bias audit before an automated hiring tool can be used, and why the vendor's own assurances do not count: the system cannot be trusted to certify itself.

When employers trusted the tool anyway

Amazon's scrapped recruiting engine (2018)

Amazon spent years building an AI tool to score applicants from one to five stars. Because it trained on ten years of resumes from a mostly male workforce, it learned to downgrade resumes that contained the word women's and graduates of two all-women colleges, as Reuters first reported. Engineers could not guarantee the system would stop finding new proxies for gender, so Amazon abandoned it.

iTutorGroup and the EEOC (2023)

The tutoring company's hiring software was configured to automatically reject women aged 55 and older and men aged 60 and older, screening out more than 200 qualified applicants in 2020. One rejected applicant was later offered an interview after resubmitting the same resume with a more recent birth date. iTutorGroup paid 365,000 dollars in the EEOC's first settlement over an AI hiring tool.

Mobley v. Workday (2025)

In a federal case in California, the court allowed a nationwide age-discrimination collective to proceed against Workday's applicant-recommendation system, covering job seekers over 40. In its own filings Workday stated that roughly 1.1 billion applications had been rejected through its tools during the relevant period, an indication of how much hiring is now decided by software the applicant never sees.

The numbers behind it

88%

of executives admit their own automated screening filters out qualified, high-skilled candidates, simply because a resume does not match the exact language in the job description.

Fuller et al., Hidden Workers: Untapped Talent, Harvard Business School and Accenture, 2021

11%

how often OpenAI's GPT ranked resumes with names distinct to Black women as the top candidate for a software-engineering role, far below every other group and enough to fail standard adverse-impact benchmarks.

Bloomberg, analysis of GPT resume ranking, 2024

The pattern in every one of these cases is the same: one automated system, one hidden ranking, and no independent check before real people were rejected at scale. That is the gap TrueStandard is built to close. Instead of trusting a single model's score, you can run the same screening judgment through four to five frontier models at once and see exactly where they disagree, in about a minute. When one model quietly downranks a candidate the others rate highly, that disagreement is the flag to review the decision by hand before it becomes a rejection.

How to verify AI hiring output

Treat any AI ranking, score, or shortlist as an unaudited draft. Before you let it filter a single candidate, run this check.

  1. 01

    Never auto-reject on a model's score alone. Keep a human in the loop for every rejection, and make sure a person can see, question, and overturn the ranking.

  2. 02

    Run a bias audit before you deploy, then repeat it. Test selection and scoring rates across race, gender, and age, and compare the impact ratio between groups. In New York City this is legally required and must be done by an independent auditor, not the vendor.

  3. 03

    Test the tool with name and demographic swaps. Submit identical resumes that differ only by name, gender cues, graduation year, or school, and check whether the ranking moves. If it does, the score is measuring the wrong thing.

  4. 04

    Check what the model is actually rewarding. Confirm it is scoring on skills and experience, not proxies like one keyword, an elite school, or an unbroken work history that screens out caregivers and career changers.

  5. 05

    Look at who gets filtered out, not just who gets through. Review a sample of the rejected pile for qualified people the tool missed, the hidden-worker problem that automated screening is known to create.

  6. 06

    Demand real evidence from the vendor. Ask for the independent audit results and validation data, and treat any marketing claim of a bias-free or fair tool as unverified until you have seen them.

  7. 07

    Give candidates notice and a human path. Tell applicants an automated tool is in use and offer a way to reach a person, which is both good practice and, in a growing number of jurisdictions, the law.

How to make AI output reliable: check it across models

The fix is not to hunt for a single more accurate model. Every large language model predicts fluent, plausible text, so each one can be confidently wrong on its own. What changes the odds is agreement. When several independent models are asked the same thing and all land on the same answer, the chance they share the exact same hallucination drops sharply. When they disagree, you have found the precise claim to check by hand before it ships.

That is what TrueStandard does: it runs your draft through four to five frontier models at once and surfaces every disagreement in about a minute, with sources. See the AI fact checker for how the method works, or read why AI cites studies that do not exist for the mechanism behind the failures on this page.

Common questions

Can I use AI to screen resumes at all?

Yes, to organize, summarize, and surface candidates, as long as a human makes the actual cut. What is not safe is letting a model auto-reject people on its score, because audits show those scores carry racial, gender, and age bias while looking perfectly objective.

Isn't an AI screener more objective than a biased human?

That is the assumption that gets employers sued. A model learns from historical hiring decisions and from name-and-demographic patterns in its training data, so it reproduces the bias in that data and applies it consistently to everyone. In the University of Washington audit it preferred white-associated names 85 percent of the time and never once favored a Black man's name over a white man's.

Is using AI for hiring even legal?

Existing anti-discrimination law still applies. The EEOC has already settled a case over an AI tool that screened out older applicants, a federal age-discrimination collective is proceeding against Workday, and New York City's Local Law 144 requires an independent bias audit and candidate notice before an automated hiring tool can be used. The tool being automated is not a defense.

What is the safest way to use AI in hiring?

Keep humans deciding, audit the tool for disparate impact before and during use, and never rely on a single model's judgment of a person. Running the same screening question through several independent models and checking where they disagree surfaces the biased or arbitrary calls that any one model, scored with total confidence, will hide.

Do not publish AI output on trust

Paste your draft. Four to five models check every claim in about 60 seconds, and you see exactly where they disagree before your name is on it.

See pricing
No Training on Your Data · 60-Second Checks · Full Verification Reports