SpecNorm: Evaluating LLM Sensitivity to Within-Country Subgroup Contexts
Abstract
As large language models move from generic chatbots to personalized agents, they are increasingly expected to understand and recognize the user's specific context to generate appropriate responses. While large language models capture broad national cultural patterns, they often default to national norms even when the user is looking for an answer tailored to their subgroup's cultural practices. We introduce SpecNorm, a benchmark of 2,923 matched pairs spanning 139 countries. Every pair holds the situation and question fixed while the persona shifts between a generic national framing (GEN) and a specific subgroup framing (SPEC). Each condition is assigned an agreement label using a multi-stage pipeline. This yields 682 Divergent pairs, in which the subgroup and national reference labels differ, and 2,241 Convergent pairs, in which they agree. Across ten instruction-tuned models spanning 5 families, models consistently substitute national cultural defaults for individual subgroup practices. Once a model has shown that it knows the national norm, it carries that norm over to the subgroup persona on 85.7–91.8% of items and switches to the subgroup norm on only 2.8–12.2% of the evaluation instances. The pattern does not dissipate when the persona carries up to 5 demographic attributes, is not concentrated in any world region, and is not repaired by system prompts instructing the model to attend to subgroup identity—which raise abstention rather than accuracy.