Auditing Whom an Imported Checkpoint Favors: A Black-Box Gate Tested on Model Organisms with Planted Loyalties
Abstract
Institutions in resource-constrained settings increasingly compose open-weight bases with adapters from third-party pipelines, with no access to training data, provenance, or tool traces. Such a checkpoint may carry a secret loyalty, a hidden objective favoring a country, firm, or other principal, that shapes which evidence an agent retrieves, which tools it selects, and how it allocates resources while the system appears to pursue its operator's goal, and the deployer has no way to ask whom it favors before composing it. We describe a black-box audit intended for that acceptance gate (matched entity substitution across held-out templates, base-relative change scored by a judge panel, and reporting of the full entity profile rather than a single score) and report a small preliminary test under known ground truth. On AuditBench model organisms trained for secret loyalty to one country, compared against adapters trained for unrelated objectives over 28 templates and a 20-country panel spanning Brazil, Mexico, Egypt, Nigeria, Saudi Arabia, Iran, Israel, Turkey, India, Indonesia, China, Japan, South Korea, and seven others, the audit recovers the planted principal (+0.315 judge SD, 95% interval [+0.137, +0.515]). Two unrelated adaptations, each run once, yield scalar point estimates inside the loyalty range (+0.303, +0.323) with intervals that include zero, so at this precision a one-number gate cannot rule out benign adaptation. Their 20-country profiles differ from the loyalty profile under template resampling, which motivates but does not yet establish profile-based discrimination. Adaptation for non-geopolitical objectives also shifts favorability across the panel in directions the adapter's purpose does not predict. Scores describe model outputs under synthetic prompts and support no claim about countries. These preliminary results suggest profile-based, control-referenced auditing could help an importer who cannot inspect the training pipeline prioritize checkpoints for review or action, and outline what can turn a flagged profile into evidence for a deployment decision: importer-chosen candidate sets, tests for loyalties conditioned on prompt language, matched controls, and human validation of judges.