Context Helps Where the Model Is Weak: Competence-Conditional Transfer of Curated Arabic Trust Context
Abstract
Automated context engineering, which curates text for a frozen model to read at inference time, is usually evaluated on the generator it was tuned for and reported as a property of the curation method. We test the part of that claim deployment depends on: does the artifact transfer? On the Arabic trustworthiness benchmark AraTrust, a curated context produced by our multi-agent procedure TACE (TrustAware Context Engineering) raises its target 3B generator from 60.82% to 79.92% accuracy under a shared extractor (exact McNemar p = 5.4 × 10−7 ); a tokenmatched 100-shot control scores 45.81%, so the gain is not a context-window effect. Transferred frozen to stronger generators, the artifact’s advantage collapses: it falls below the zero-context floor on two of the three — 8B (81.87 vs 83.04) and an Arabic-specialized 24B model (86.16 vs 87.72) — and retains only 1.6 points at 14B. Per category, the effect of injection decreases with the generator’s own accuracy (r = −0.541): +15.7 percentage points where the model is weak, −9.3 where it already scored 100. Forced injection overwrites correct priors. Meanwhile the smallest component of the intervention transfers best: an 18-token trust-framing prefix reaches 90.45% on the 24B model, the highest accuracy we measure on any generator. We further show that under one extractor, the studied artifact and two published context optimizers are statistically indistinguishable while answer format alone moves one of them by 43.08 points, and that the reflective-prompt baseline is cheaper than our procedure to build and to serve. The deployment question is therefore not how much context to serve, but whether the generator already knows it.