Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
Abstract
Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Six open-weight language models are evaluated across these prompts, producing 12,600 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji-augmented and invisible Unicode variations score 16.89\% and 14.33\% near baseline, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.28\%, 0.11\%, and 2.00\%. Inspection of raw model outputs reveals four recurring response behaviors: hallucinated benignity, structural collapse, language drift, and null responses.