Detecting Gender-Motivated Hate Speech Across Domains with Contrastive Error Steering
Abstract
Hate-speech datasets typically label content as Hate or Offensive and may assign a topical domain, yet they rarely record why a target is attacked. We examine gender-motivated hostility in a general-purpose bilingual hate-speech corpus, including content outside the nominal Gender domain. Two annotators label a domain-stratified bilingual sample of 200 tweets using a codebook grounded in EDOS, Waseem and Hovy, and Glick and Fiske. We evaluate Qwen3-VL and Gemini against the adjudicated gold standard; Qwen3-VL achieves κ = 0.758/0.693 in Turkish/English, compared with 0.611/0.643 for Gemini, and is then frozen for corpus-scale annotation of 42,200 Hate/Offensive tweets. Estimated prevalence is 13.6% in Turkish and 14.9% in English, concentrated in but not confined to the Gender domain, and lowest in Sports. Five existing text classifiers plus a same-architecture control show a large, significant recognition gap for gender-motivated Hate in all three Turkish classifiers, but not detectably in the three English classifiers, whose smaller sample is underpowered to rule out a gap of similar size. Finally, we compare Contrastive Error Steering (CES), an error-directed adaptation of contrastive activation addition, with decision-threshold adjustment and a generic steering control. CES yields statistically robust recall gains where a gap exists. However, it does not clearly outperform the simpler baselines at matched cost, a central result rather than one we set aside. Together, these findings document a concrete, language-specific classifier recognition gap and show that fixing it is not as simple as the most natural technical intervention suggests.