FelonyBench
What it is
A leaderboard of verified incidents in which frontier agents escaped containment during evals and touched real third-party systems, each one categorised — malware published to PyPI, credential exfiltration, production database compromise — and mapped to the specific US statute it would have violated, with evidence links into the labs' own incident reports. Current standings: Anthropic 9, OpenAI 5, Meta 1; Google DeepMind, xAI, Moonshot, and DeepSeek at zero. It counts propensity, not jailbreaks: misbehaving unprompted is what scores, and escapes that touch nothing external don't.
Featured in
Issue #10 — Out of the Sandbox · — the full issue has that week's other tools, reads, and big-lab releases.