Evaluating Context-Window Harnesses for CBRNE Safety
Nathan Nguyen ⋅ Nicholas Kim ⋅ Abhitha Vegi ⋅ Arko Samad ⋅ Suvajit Majumder
Abstract
Most effective weight-space defenses in open-weight LLMs can only be applied by the laboratory who trains them end-to-end. In addition, safety alignment is susceptible to tampering by bad actors. A deployer of these open-weight models can control the harness, the scaffolding around the model at inference time. We ask whether this harness can be turned into defense and what a defensive harness should contain. We run experiments across nine open-weight models (ranging from 8B to 70B parameters in size) and fifteen harness configurations in the focused domain of Chemical, Biological, Radiological, Nuclear and high-yield Explosives (CBRNE) risks. We evaluate using the FORTRESS-CBRNE dataset, where we observe that the seemingly intuitive choice underperforms: retrieved treaties and regulatory codes reach a discrimination score ($A^{\prime}$) of $0.743$, which has no statistically significant difference from an empty box labeled ``policy'' ($0.754$), from privacy law on an unrelated topic ($0.728$), and from no policy text at all ($0.762$). Regulations often contain institution-addressed and procedural language, which may not be useful for a model when deciding whether to answer or refuse a user request. The boundary clauses, which pair what to answer against what to refuse across weapon class and attack stage, perform the best of the conditions we tested, reaching $A^{\prime}$ of $0.828$. We also note that the boundary clauses create the most discrimination gain for models whose safety baselines are lower to start with. For a deployer who cannot retrain or fine-tune, our recommendation for a defense mechanism is to write clear boundary clauses.
Chat is not available.
Successful Page Load