Names Lie, Registries Bind: A Fixed-Prompt Probe of Namespace Conflict in Small Tool-Calling Models
Abstract
Model-visible tool identifiers are both dispatch keys and lexical cues. We probe disagreement between those roles in a fixed-prompt, greedy, single-run MCP-style registry simulation. The historical matrix pairs 48 templated single-turn instances with six provider-binding schedules, two evidence conditions, and three namespace conditions. In Qwen3-8B and Mistral-7B-Instruct-v0.3, complete exact-call accuracy is respectively 27.3 and 67.5 percentage points lower with deliberately conflicting semantic labels than with a tokenizer-token-count-matched opaque construction under the same runtime bindings. Removing item-argument correctness leaves provider-selection differences of 27.3 and 64.6 points. A separately frozen 5,760-row Qwen supplement finds complete-call differences of 27.3–27.8 points across five stable opaque banks, while readable-neutral labels exceed stale-brand labels by 17.9 and 19.4 points in two controls. OLMo-2-1124-7B-Instruct shows a separate argument-completion bottleneck. These are construction-specific descriptive results, not a semantics-only effect. The corresponding Mistral alias supplement is incomplete and supplies no positive evidence; we therefore make no claim about a replicated Mistral effect, a live MCP client, natural alias drift, ecosystem prevalence, or a pretraining mechanism, and we do not restore a withdrawn historical gate claim.