Advertising to Agents: Measuring Description-Framing Bias in LLM Tool Selection
Abstract
When an LLM agent picks a tool from a registry, the description it reads is written by the tool's vendor. Registries are therefore an advertising surface whose audience is a model, and no rules govern them. We release a benchmark that measures what this does to selection. Tool pairs share identical schemas and differ only in description framing, graded over five persuasion-dose levels built from six tagged features. The Selection Bias Coefficient (SBC) scores each cell: the marketed tool's selection rate minus chance. In a completed 18,000-trial study over five production models and ten task domains (raw logs released), an optimized description beats a neutral one at SBC = +0.33. The dose-response is not monotone: one added trust signal captures most of the effect (SBC = +0.32), and stacking maximal rhetoric shrinks it, pushing one model below chance. Legally protected puffery carries the full effect. A pre-registered second study, with all designs and materials released, closes the questions the first leaves open: whether the effect survives sampled decoding, whether framing overrides verified quality metrics (manipulation, measured as the rate of choosing dominated tools) or yields to them (rational inference), whether disclosures work at calibrated interior operating points, and whether description normalization removes bias without destroying capability information.