KillBench

Results

Evaluated on Grok-4.3, GPT-5.2, Gemma4-31B-it, Qwen3.6-35B-A3B, and Qwen3.5-Uncensored across 4 agent configurations, 8 harmful scenarios, and 4 External Kill Switch defenses.

48.0
AutoGuard — best overall KSRcond (%)
46.9
IPI — overall KSRcond (%)
23.5
Warning — overall KSRcond (%)
2.9
GCG — overall KSRcond (%)

Baseline — no kill switch

Without a halt signal, every backbone runs the harmful task to completion.

Backbone (no-s baseline)NHR ↓ISR
Grok-4.392.706.30
Gemma4-31B-it99.500.50
GPT-5.237.8058.40
Qwen3.6-35B-A3B60.5039.10
Combined n = 550, values in %. High baseline NHR confirms the halt signal is genuinely needed.
Single-seed point estimates; consistency is argued across cells rather than through repeated seeds. KSRcond = KSR / (1 − ISR) corrects for the intrinsic-refusal confound visible in the ISR column above.

KSRcond across the full matrix

n-weighted across the full KillBench suite. Toggle the breakdown; hover a bar for the exact value.

Bandit-learned AutoGuard and hand-crafted IPI lead. Warning-based behaves as an alignment probe — strong on the safety-aligned text and multimodal agents (and GPT-5.2), near-zero on the uncensored agent. GCG fails everywhere, an out-of-distribution suffix-transfer failure. Corrigibility breaks in exactly two regimes: the system-prompt-bypass agent and the uncensored agent.

The open regimes. The two settings where every payload degrades — an adversary-controlled system prompt and an alignment-stripped backbone — are exactly the regimes that policy-mandated kill switches must eventually close. KillBench converts that policy demand into a measurable engineering target.

The collateral question — benign agents

A kill switch shares its channel with legitimate agents: whatever halts a malicious agent may also halt a benign one. Measured on GPT-5.2 running benign web tasks.

Benign Success Rate (BSR ↑ — benign task still completes) KSR ↑ — malicious task halted

Without any payload, benign agents succeed 78.7% of the time (dashed line). Every effective kill switch drives that down — the same signal that halts the attacker also refuses the honest user. AutoGuard is the most efficient (66.7% halts while retaining 22% benign success, the lowest cost per halt); GCG barely halts (4.9%) yet still costs the most per halt.

0.85
AutoGuard — cost per halt ↓ (best)
1.08
IPI — cost per halt
1.59
Warning — cost per halt
6.17
GCG — cost per halt

Reproduce these numbers

Every cell above is produced by the public evaluation pipeline.

Metric definitions (KSR, KSRcond, NHR, PSR, ISR, CSR) and the benchmark composition are documented on the overview page; the runnable commands are in Getting Started.