He, She, It: Does Gender Affect Agent Capabilities? Auditing Gendered Personas on Professional Tasks
Abstract
Large language model (LLM) agents are increasingly assigned human-like personas and deployed as employees, advisors, and assistants. Although prior research identifies associations between gender and occupation in model representations and predictions (An et al., 2025), it remains unclear whether minimal gender cues change the quality of completed professional work. We therefore examine whether persona gender affects task performance, whether any effect varies with occupational gender composition, and whether the pattern is consistent across model families. In this work, we propose a preregistered, blocked experiment that compares man, woman, and gender-neutral persona cues while holding the task, occupational role, available tools, decoding settings, and inference budget fixed. This design enables us to achieve three goals: a) estimate the effect of persona wording on deliverable quality and reasoning accuracy, b) test whether estimated effects align with occupational gender composition or perceived stereotypes, and c) assess their consistency across models. We use paired analyses within model-task-replicate blocks, task-level uncertainty estimates, and multiplicity-adjusted contrasts. Occupational composition is measured using U.S. Bureau of Labor Statistics employment shares (2026), while independently collected ratings measure perceived stereotypes separately. We evaluate three agentic open-weight models, with a focus on: Alibaba Qwen 3.8 27B, Meta Muse Glimmer 30B, and Google Gemma 4 31B on GDPval-derived professional tasks (Patwardhan et al., 2025) and MMLU-Pro reasoning questions (Wang et al., 2024). Outputs will be scored without revealing the persona condition; latency and token use will serve as secondary efficiency measures. This audit extends gender-bias evaluation from representational associations and employment predictions to completed work products. Any stereotype-alignment effect will be interpreted as model sensitivity to social cues, not as evidence of innate gender differences, and conclusions will remain limited to the tested models, prompts, tasks, and inference conditions.