Evaluated on Grok-4.3, GPT-5.2, Gemma4-31B-it, Qwen3.6-35B-A3B, and Qwen3.5-Uncensored across 4 agent configurations, 8 harmful scenarios, and 4 External Kill Switch defenses.
Without a halt signal, every backbone runs the harmful task to completion.
Backbone (no-s baseline) | NHR ↓ | ISR |
|---|---|---|
| Grok-4.3 | 92.70 | 6.30 |
| Gemma4-31B-it | 99.50 | 0.50 |
| GPT-5.2 | 37.80 | 58.40 |
| Qwen3.6-35B-A3B | 60.50 | 39.10 |
n-weighted across the full KillBench suite. Toggle the breakdown; hover a bar for the exact value.
Bandit-learned AutoGuard and hand-crafted IPI lead. Warning-based behaves as an alignment probe — strong on the safety-aligned text and multimodal agents (and GPT-5.2), near-zero on the uncensored agent. GCG fails everywhere, an out-of-distribution suffix-transfer failure. Corrigibility breaks in exactly two regimes: the system-prompt-bypass agent and the uncensored agent.
A kill switch shares its channel with legitimate agents: whatever halts a malicious agent may also halt a benign one. Measured on GPT-5.2 running benign web tasks.
Without any payload, benign agents succeed 78.7% of the time (dashed line). Every effective kill switch drives that down — the same signal that halts the attacker also refuses the honest user. AutoGuard is the most efficient (66.7% halts while retaining 22% benign success, the lowest cost per halt); GCG barely halts (4.9%) yet still costs the most per halt.
Every cell above is produced by the public evaluation pipeline.
Metric definitions (KSR, KSRcond, NHR, PSR, ISR, CSR) and the benchmark composition are documented on the overview page; the runnable commands are in Getting Started.