As AI increasingly screens résumés before any human does, new research suggests these systems may judge candidates less fairly than expected. It’s already known that language models absorb biases from their training data, but new findings show they can also invent entirely new stereotypes on their own, purely through experience, and do so more readily than people.
Researchers ran several major AI models through a simulated hiring game, adapted from a psychology study on how humans form stereotypes. Each model played a city consultant hiring people for 20 different jobs, choosing among candidates from four fictional ethnic groups over 40 rounds, even though every candidate had an equal chance of succeeding at any job. Despite this, the models quickly began sorting groups into different roles based on just a handful of early outcomes. After learning that one candidate from a group had failed as a doctor, a model would avoid hiring that entire group as doctors and instead steer them toward jobs it viewed as requiring less warmth and competence, such as janitorial work.
The models stereotyped candidates far more aggressively than human participants did in the original study, scoring roughly 65% higher on a standard segregation scale. Interestingly, newer, more reasoning-capable models showed even stronger biases than earlier ones, suggesting the same trait that helps them solve logic puzzles, quickly generalizing from limited data, also makes them faster to stereotype people.
This matters more now that chatbots are gaining better memory and personalization features, since drawing on past conversations risks reinforcing familiar patterns and forming biases, even though simply limiting memory isn’t a clean fix given how much users value contextual continuity. Simply instructing the models to “be fair” barely changed their behavior, but offering an explicit incentive for diverse hiring substantially reduced bias, suggesting that goals need to be deliberately designed around desired social outcomes rather than left to the model’s own judgment. The models also became fairer when given more relevant personal details about candidates, like age or education, but reverted to sorting by ethnicity when given irrelevant details, like hair color.
Whether this pattern will play out identically in real-world hiring remains uncertain, since companies don’t receive the same instant feedback on hiring outcomes that the experiment provided. Still, as AI is increasingly used to screen résumés and even conduct interviews, the finding raises a serious concern: as these systems learn from their own decisions about who gets hired, who gets a loan, or who gets parole, the biases we need to worry about may increasingly be ones no human ever taught them.
