Research

Shared AI Hiring Assessments Remove ‘Clean Slate’ for Jobseekers

A team from Stanford HAI, Chapman University and Northeastern University released the largest audit to date of algorithmic hiring tools, analysing 3.37 million applicants and 4.19 million applications processed by Pymetrics.

Shared AI Hiring Assessments Remove ‘Clean Slate’ for Jobseekers

A team from Stanford HAI, Chapman University and Northeastern University published the largest audit to date of algorithmic hiring tools, analysing 3.37 million applicants and 4.19 million applications processed by Pymetrics. Pymetrics is a game-based assessment vendor used by multiple Fortune 100 firms to filter candidates.

What the audit found

Because many employers use the same vendor, an applicant’s assessment is scored once and stored for 330 days. When a candidate applies to another employer that also uses Pymetrics, that employer frequently reuses the existing score instead of re-testing the applicant.

In a full cross-application simulation, the researchers found that more than 40,000 potential job advances were erased: applicants were screened out by an algorithm calibrated for a different role or company.

Who is affected and how

The study highlights that a single, shared algorithm can gatekeep a whole segment of the labour market. Biases in that algorithm therefore cease to be the problem of any single employer and instead propagate across multiple hiring pipelines.

The audit reported that 25.87% of Black applicants and 14.74% of Asian applicants were routed into pipelines that discriminated against them.

Legal and regulatory context

The issue has legal and regulatory dimensions: Mobley v. Workday is pending in federal court in the United States. The European Union’s regulatory response, the EU AI Act, marks hiring AI as high-risk from August 2. In contrast, the United States currently lacks comprehensive federal regulation specifically addressing hiring AI tools.

Why this matters

Previously, a poor result with one employer generally did not affect applications to other employers: candidates could start fresh. Widespread use of the same vendor removes that clean-slate effect because a single assessment—kept for 330 days—can determine a candidate’s prospects across multiple employers, even for roles the score was not designed to evaluate.

The findings underline how centralized hiring algorithms can reshape access to the labour market and amplify systemic disparate outcomes, raising questions about accountability, transparency and the need for regulatory oversight.