Audit Before You Inject: Gains, Costs, Transfer, and Defects of a Curated Prompt Context on AraTrust
Abstract
Small language models increasingly serve as the executors of agentic systems, and the teams shipping them rarely get to fine-tune, so the adaptation that reaches production is the cheapest one: prepend curated guidance to the prompt — often guidance a larger model wrote. We audit one such artifact end to end — gain, cost, transfer, defects. On the Arabic trustworthiness benchmark AraTrust (n=171), a multi-agent-curated Arabic context raises its target 3B generator from 60.82 to 79.92 (exact McNemar p=5.4e-7); a token-matched 100-shot control scores 45.81, so the gain is not a token-budget effect. Transferred frozen to three larger generators, the artifact falls below the zero-context floor on two. On Mistral-Saba-24B, an 18-token framing-and-format prefix is not less accurate on observed means than the corpus (90.45 vs 86.16) while serving about 50x fewer tokens (means 113.7 vs 5,774) and 107 ms faster per call; on the 3B the same prefix is 13.25 points worse. Build cost is 41x a reflective-prompt baseline's, per-call tokens nearly 8x, so serving does not amortise the build. A proposed competence condition — injection compresses per-category accuracy toward the mean (slope 0.452) — is directionally robust but not significant under clustering-aware inference (p=0.21–0.22). The firmest findings concern measurement and provenance: answer format alone moves one baseline 43.08 points; the artifact embeds an answer sheet with one fabricated entry served to 153 of 171 test items; two of twenty items in the benchmark's weakest category are mislabelled. We release the artifact as deployed and repaired, per-item predictions, and scripts reproducing every number.