Researchers from the University of Pennsylvania and the World Bank examined how well large language models (LLMs) predict human reactions to social norm violations. Their paper, posted on arXiv, is based on a new dataset called NormReact, which contains 450 norm-violation scenarios. The team measured how closely LLM estimates of emotional and behavioral responses align with those given by human respondents.
The scenarios recorded the violator’s gender and the relationship between the violator and observers (for example, close relative, acquaintance, or stranger). For each scenario the study posed 16 questions covering the violator’s and bystanders’ emotions, intensity of those emotions, expected behaviors (public shaming, gossip, avoidance, confrontation, or inaction), and what reaction would be considered socially appropriate. Human participants answered the same questions, allowing a direct comparison between human expectations and model predictions.
Models tested
The study evaluated six large language models: GPT-5.2, Claude-4.5-Opus, Gemini-3-Pro, GPT-OSS-20B, Llama-4-Scout-17B and Gemma-3-12B. Each model received the same 16-question survey across the 450 scenarios.
Main findings
All six models showed the same pattern: they systematically overestimated the likelihood of negative social consequences—such as disapproval, gossip, avoidance or confrontation—compared with human respondents. The discrepancy was most pronounced in scenarios where human respondents expected inaction; models frequently predicted punitive or condemnatory responses even in these cases.
According to the authors, current LLMs portray a "harsher social world": they attribute negative sanctions to situations where people are more likely to be passive, tolerant, or restrained. Agreement between human and model responses also decreased as social distance between actors increased (for example, differences were larger for reactions by strangers).
Why this matters
These distortions have practical implications. If LLMs are used for conflict mediation, modeling social processes, or policy simulations, their bias toward predicting punishment could lead to misleading conclusions. Systems might overemphasize social condemnation and understate tolerance or context-dependent restraint.
The researchers argue that developers must tackle a new task: AI should not only learn which rules people regard as important, but also what actually happens when those rules are broken—who reacts, how strongly, and how often.
Related risks and a military context
The paper’s findings are particularly relevant given rapid development of battlefield-capable, autonomous systems in countries such as China and the United States. The article notes earlier reporting that the People’s Liberation Army and other actors are accelerating transfer of advanced technologies to military testing grounds. Technical and ethical obstacles remain: the International Committee of the Red Cross (ICRC) has warned that autonomous weapons may have difficulty reliably interpreting signals of surrender, since doing so often requires evaluating context, behavior and intent together.
The study’s authors point out that similarly complex pattern recognition is required for interpreting human social responses to norm violations, and current AI systems are prone to misreading these patterns.
Conclusions
The results indicate that current LLMs tend to overpredict punitive and condemnatory reactions to social norm violations compared with human expectations. This bias could distort outcomes in applications that require accurate modeling of human behavior, so developers should prioritize improving models’ understanding of metanorms—how people actually respond to transgressions in varying social contexts—if they want more reliable, human-aligned systems.



