Researchers have found that language models are capable of independently forming new social prejudices
7/22/2026, 12:55 PM • Евгения Слив

Researchers from the Universities of Princeton and Chicago have published a paper that changes the understanding of the nature of bias in artificial intelligence. Scientists have proved that modern language models are capable of independently forming new social stereotypes, even if there was no discrimination in the training data initially. To test this hypothesis, a classic recruitment experiment was replicated, in which participants distributed candidates by vacancy and received feedback on the results. All the candidates belonged to one of four fictional ethnic groups that had no real differences and had equal chances of success. While people developed only moderate biases based on negative experiences over time, language models showed a significantly higher level of stratification.
The identified problem is based on the fundamental principle of the interaction of algorithms with data, known as the balance between research and use. When making a decision, the system chooses between testing unknown options and relying on previously successful strategies. The researchers tested fifteen models from leading technology companies and found a clear pattern: the newer and more powerful the algorithm, the stronger the bias. Advanced models draw more accurate conclusions from past results, preferring candidates from a group that has previously been associated with success. This behavior, which seems rational from the point of view of maximizing rewards, turns into a systemic dysfunction, since the model completely stops exploring alternatives and thereby marginalizes entire categories.
The practical implications of this discovery are already evident in the corporate sector, where more than ninety percent of companies use algorithms in their recruitment processes. The use of automated systems in personnel management has become the subject of close attention from regulators and courts. Major HR management software vendors are facing lawsuits accusing their tools of discriminating against candidates. Similar claims have been made against technology giants, whose systems allegedly infringed on the rights of employees during organizational changes. The task of developers now is to create mechanisms that selectively suppress harmful pattern matching, while maintaining constructive forms of abstraction that make language models effective tools.
