Hiring AI Promised Objectivity. It Delivered a Single Point of Failure.
The largest-ever audit of AI hiring shows what happens when one vendor's score becomes everyone's verdict. The cost efficiency may be quietly destroying both objectivity & competitive edge.
A continuing thread. This is the second piece in a line of thinking about evaluation efficiency. In the first — on AI and testing — I argued that when we make assessment cheap and scalable, we quietly strip out the things that made it trustworthy: independent judgment, validity, and the second chance. What follows is the same pattern wearing a different suit. The domain is hiring; the mechanism is identical.
Read part one: “AI Can Grade a Million Essays Overnight. That Doesn’t Fix Testing”
Researchers at Stanford HAI, Chapman, and Northeastern just published the largest audit of AI hiring algorithms ever run — 3.37 million applicants, 4.19 million applications, all funneled through a single game-based assessment vendor that Fortune 100 firms buy to filter candidates.
Here’s the mechanic that matters. You apply, you play the assessment, you’re scored once — and that score is stored for about 330 days. Because so many employers run the same vendor, the next one doesn’t re-test you. It pulls the old score. A cross-application simulation estimated that this quietly erased more than 40,000 job advances: people knocked out by an algorithm calibrated for a different role, at a different company, months earlier.
Most of the coverage will be about bias, and the bias is real — the audit found Black and Asian applicants disproportionately routed into pipelines that disadvantaged them. But I want to set bias aside deliberately, because it’s the payload, not the mechanism. A perfectly unbiased version of this system would still be dangerous, and the reason it’s dangerous is the part nobody is pricing: what cost efficiency actually costs you.
The thing you bought was not what you think you bought
Firms adopt these tools to buy objectivity — a clean, standardized, data-driven screen that strips out messy human judgment. That’s the pitch, and it’s seductive.
What they actually get is the opposite. A single vendor’s model, trained and calibrated for one set of roles, produces one score, and that score is then reused across companies and jobs it was never validated for. You didn’t replace subjective judgment with objective measurement. You replaced your judgment with someone else’s, froze it, and relabeled it “objective.”
There’s a precise word for what the market now has, and it isn’t objectivity. It’s uniformity. Objectivity would mean each employer independently measures a candidate against its own validated criteria and the scores converge because they’re all tracking something real. Uniformity means everyone gets the same answer because everyone is asking the same oracle. Those look identical from the outside — wide agreement! — but they are opposites. Agreement produced by a shared instrument isn’t corroboration. It’s an echo. And an echo tells you nothing about whether the original sound was right.
Independence was load-bearing, and it’s gone
The labor market quietly relied on a feature it never named: forgetting.
When every employer evaluated you fresh, the errors in those evaluations were uncorrelated. A bad day, a bad fit, a bad rater at one firm got washed out at the next. Independent looks gave a candidate many shots at being seen correctly, and gave the market many shots at catching a good person the last screen missed. That redundancy felt inefficient. It was actually the error-correction.
Caching plus vendor concentration removes the forgetting. Now one assessment of you determines many outcomes at once. The fifty independent coin-flips become one coin, flipped once, photocopied fifty times. The variance that used to protect candidates — and protect employers from their own screening mistakes — is gone. Not because the model is biased, but because it’s shared. The efficiency gain and the loss of second chances are not two things. They are the same act.
Now the part that should keep executives up at night
Set the candidate’s harm aside and look at this purely as a competitive question, because this is where the cost story inverts.
If you and all your competitors screen through the same vendor, pulling from the same cached scores, then you are all fishing in the same pre-filtered pond. You surface the same “top” candidates. You reject the same people. In lockstep. Your talent funnel becomes a near-perfect copy of your rivals’ funnel — and you paid a subscription fee for the privilege of making it so.
But competitive advantage in hiring has only ever come from one thing: seeing value the rest of the market misses. The great hire is the person everyone else screened out who turns out to be excellent — the mispriced candidate. Edge in a talent market lives in disagreement. It is, by definition, the refusal to converge on the consensus view of who’s good.
A shared screen erases disagreement by construction. If the whole market runs the same model, there is no mispricing left to exploit, because the market has agreed in advance — through a single algorithm — on exactly who is and isn’t worth looking at. You have bought a product whose core function is to eliminate the only source of talent alpha there is.
And here’s the kicker for the firm willing to think two moves ahead. Those 40,000 erased advances aren’t only harmed applicants. They are universally invisible candidates — people the entire market, including your competitors, can no longer see. Which means the firm that opts out, screens independently, and looks at that rejected pool gets first pick of strong candidates at a discount, precisely because everyone else has agreed not to look. Independence stopped being an ethics nicety. It became an arbitrage.
So tally the real trade. You saved a few dollars per applicant on screening — a rounding error against the cost of one bad hire or one missed great one. In exchange you gave up your objectivity (you now inherit a vendor’s calibration as if it were neutral truth), your independence (you no longer make your own call), and your differentiation (your workforce is sourced from the same filtered pond as everyone else’s). The efficient firm and the competitive firm have quietly stopped being the same firm.
Before anyone calls this overblown
A fair critique has to survive the obvious rebuttals, so let me make them myself.
“Candidates were never independent — references, shared ATS data, and back-channels already correlated outcomes.”True. The honest claim isn’t that we lost a pristine clean slate. It’s that we industrialized and automated a correlation that used to be partial, leaky, and slow. That’s worse, not better, but the argument has to be precise to land.
“Reuse isn’t intrinsically bad.” Also true — stable, valid, portable signals help people; that’s the logic of licensure and standardized credentials. The problem here is unaccountable, context-blind, concentrated reuse: a score the candidate can’t see, contest, or refresh, validated for one context and applied to another, behind a vendor common enough to be inescapable.
“The 40,000 is a simulation.” Correct, and worth saying plainly. It’s a modeled counterfactual, not a confirmed body count. It tells you the scale of the exposure the mechanism creates — which is the right thing to worry about — not a tally of proven victims.
And the deepest one, which the headlines miss: is the game even valid? If a behavioral game’s link to actual job performance is weak, then caching isn’t freezing a fair score — it’s freezing noise and handing it the authority of a verdict across an entire market. Portability doesn’t just spread bias. It can make a possibly-invalid instrument load-bearing for millions of careers.
The regulators are circling the wrong target (so far)
Mobley v. Workday is in federal court testing whether a vendor can be liable as the employer’s “agent” under anti-discrimination law — which is the right pressure point, because the vendor is the concentration. The EU AI Act flags hiring AI as high-risk. The US has essentially nothing. But notice all of this aims at bias. None of it yet addresses the architecture — the caching, the concentration, the cross-context reuse — which would be harmful even with bias fully solved.
The bottom line
The market got more efficient by learning to remember. But forgetting was the feature. Independent, redundant, slightly wasteful evaluation was what gave people second chances and gave firms a shot at the candidates their rivals fumbled. Strip it out and you don’t get a fairer, sharper market — you get one verdict, everywhere, that no one can appeal and no one can out-compete.
The clean slate used to be free. Now it expires in 330 days, set by a test you forgot taking, for a job you never got.
And if your screening funnel is now identical to your competitors’, it’s worth asking where, exactly, your talent advantage is supposed to come from.


