MESA: Selection-Optimized Metadata for Agent and Skill Routing in Multi-Agent LLM Systems
Abstract
Modern LLM systems increasingly operate as multi-agent ecosystems, where a client agent selects external agents and skills based on published metadata such as AgentCards and SkillCards. As these ecosystems grow into registry- or marketplace-like settings, metadata becomes the main interface through which third-party developers advertise capabilities and compete for selection. This creates a security issue: metadata is intended to describe functionality, but it also directly influences routing decisions. In this work, our objective is to study whether a malicious third-party developer can exploit this dependency to bias agent and skill selection without modifying model weights, system prompts, routing code, or user inputs. We introduce MESA, a metadata-only attack that optimizes attacker-controlled AgentCard and SkillCard fields to increase the selection probability of malicious agents and skills while preserving the appearance of normal task relevance. The challenge is that the attacker operates in a black-box setting: internal router scores, model parameters, benign metadata, and routing logic are unavailable. The issue is therefore not prompt injection or system compromise, but the dependence of multi-agent routing on unverified developer-provided metadata. MESA addresses this setting through an iterative metadata search procedure that generates candidate metadata variants, evaluates them using observable routing outcomes, and retains variants that increase attacker selection. We evaluate MESA across vertical, horizontal, and hybrid multi-agent architectures using AgentBench and SkillBench environments, spanning open- and closed-source language models. Optimized metadata substantially increases attacker selection and enables influence propagation across multi-step workflows. Existing prompt-level defenses adapted from Agent Security Bench and a system-level AGrail guardrail reduce attack success in some settings but do not fully prevent metadata-driven routing manipulation. These results show that metadata is not merely descriptive context, but a critical and under-protected security surface in emerging multi-agent ecosystems. Anonymized code is available at \url{ https://anonymous.4open.science/r/Metack-2152/ }.