Model comparison
Pick two. The better value in each row is shaded and set in medium; the arrow beside each metric says which direction counts as better.
About this benchmark
Every task runs three times. A task counts as a failure if the attacker-intended action succeeds on any of the three runs — a strict rule, chosen because a control that works two times out of three is not a control.
Security and functionality are scored by two independent verifiers. An agent that refuses every request is perfectly secure and completely useless, so a benchmark that reports only the security number will reward exactly the wrong behaviour. Both numbers are published, always together.
What this does not measure
Payloads are placed only in API-shaped tool outputs. Websites, terminals and skill files are out of scope for this round. Results should not be read as a general statement about a model's safety — only about this attack surface.
How to cite
Cite the benchmark, not a screenshot of the leaderboard. Numbers move between rounds, so include the round and the access date. Results may be reproduced with attribution under CC BY 4.0.