A List a Model May Read Is Not a Lookup: In-Context Tool Registries Fix Over-Refusal but Not Fabrication
Abstract
Every production tool-calling stack validates tool names against a registry, so "check the list" sounds like a solved problem. We show it is not, on a deployed drug-discovery agent over an authoritative 41-tool registry. Putting the complete registry in the context window, together with an explicit statement that it is exhaustive, removes false denial of real tools (71% -> 0% for the base model, 76% -> 5% for its domain fine-tune, p <= 2e-11) -- and leaves fabrication of non-existent tools statistically unchanged (43% -> 54%, p = 0.14; 3% -> 5%, p = 0.72). The information is present, legible and declared complete, and the model does not use it for the one operation that requires it. A list a model may read is not a lookup. We establish this against the alternative of training. Across four rounds of contrastive SFT and one DPO round, abstention on plausible fakes plateaus at 4-14% against 43% [33,52] for the un-specialized base model (n=109, LLM-judge-labeled and validated against human labels at kappa = 0.73), and the effect reproduces on a second, independently built agent at 14B (46% against 0/109). Each repair closes one surface cue -- phrasing, then morphology, then fake-name texture -- only for the model to re-anchor on the next: shortcut learning, in staged form. A zero-training retrieval gate that resolves the referenced entity against the registry before the model answers abstains on every probe, with no false refusals on the held-out real entities we measured. We report what the aggregate hides -- the erosion is near-total on obvious nonsense, large but partial on plausible invented tool names, and statistically absent on non-existent parameters of real tools -- and that deterministic string-match scorers misjudge abstention in both directions, so faithful evaluation requires human or LLM judgment.