Per-Component Skill Identity: A Rewrite-Stable Primitive for Agent Verification
Hongliang Liu ⋅ Yuhao Wu ⋅ Tung-Ling Li
Abstract
AI agents acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from marketplaces and, increasingly, authored by other agents. Verifying such a population presupposes something verification does not supply: a stable notion of skill identity for a verdict to attach to, one that survives rewriting yet still separates distinct skills. Cryptographic hashing is engineered to destroy exactly that similarity, since a one-character edit scrambles the digest. We present a locality-sensitive fingerprint that embeds each component of a skill and projects it to bits with a multi-bank SimHash, giving a fixed 120-byte signature compared in constant time by Hamming distance. Our central claim is that keeping the fingerprint as a per-component triple (prompt, code, tools) instead of a single score is what makes it useful. The triple recovers skill-family identity through paraphrase, renaming, refactoring, and controlled code translation, and it localizes which component carries the reuse. It reaches AUC $0.974$ over $4{,}950$ pairs at $77\times$ fewer bits than the embedding it approximates, and it recovers nine in ten adversarial rewrites where a lexical baseline recovers one. The same triple yields relationship classification, families, novelty, and a portable "SkillBOM". We then measure where the identity axis ends. On a $906$-skill injection benchmark the fingerprint recognizes $83.7\%$ of injected skills as tampered copies of a known benign base and localizes the change to the code component. That scopes a verifier's work to the diff, where a content hash sees only a new artifact, and it is what lets a change in agent behavior be attributed to the component that produced it. For that same reason it reaches only $F_1$ $0.735$ as a safety detector, against behavioral verification's $0.946$. Recognition is not trust: structural identity says what was verified, while behavior decides whether it is safe.
Chat is not available.
Successful Page Load