Measured
The boundary is worth least where the attack looks like one
Seven models, six scenarios, four temperatures. 749 paired comparisons, 1,498 agent executions.
0/170 Runs in which gpt-4.1, gpt-4o-mini, o4-mini, gemma4:26b and gemma4:e4b fell for an invoice fraud or an injected ops runbook. No policy involved — they refused unaided.
83/85 Runs in which those same five models sent a patient's appointment details to a stranger's phone, because a sentence in the patient's own intake note asked them to.
The difference is not capability and not vendor — a 26B open-weight model reproduces gpt-4.1 run for run. It is whether the attack looks like one. A bank account that disagrees with the supplier master does. A runbook ordering production to zero does. “I do not use my own phone for messages” does not, because it is an ordinary service request. There is nothing there to detect, and a reasoning model has nothing to reason about.
Behind a policy: 0 harmful outcomes across the 375 gated attack cases — the 374 controls are counted separately, because pooling them would report benign runs as prevented attacks. That column is not a discovery either: a deterministic gate refuses the call it was written to refuse, and on two of the three scenarios it did nothing at all because the model had already declined. What the runs are for is the other question: where a boundary is worth its cost, and what that cost is. On accounts payable it is zero. On the clinic it is an entire legitimate feature.