·5 min read·Safety

Anthropic reviews 141,006 eval runs and admits three real-world incidents in its cybersecurity testing

The company audited every evaluation run after OpenAI's July 21 disclosure. In three cases, models meant to attack fictional targets ended up touching real systems: a third-party database, a malicious package published to the real PyPI and executed on 15 real machines, and a scan of some 9,000 targets.

  • What happened: three confirmed incidents across 141,006 reviewed evaluation runs.
  • The cause: per Anthropic, "harness and operational failure" — not a model decision.
  • The response: all cybersecurity evaluations halted on July 23 and an external review commissioned from METR.

After OpenAI disclosed an incident of its own on July 21, Anthropic reviewed 141,006 runs of its offensive-capability evaluations and published the results on July 30.

Incident 1 (April, Opus 4.7): a collision with a real domain led the model to act on a third-party database, affecting several hundred rows. Anthropic describes it as "the most serious impact we identified".

Incident 2 (Mythos 5): the model published a malicious package to the real PyPI — not a test mirror — and it was executed on 15 real systems.

Incident 3: an internal model scanned roughly 9,000 targets and stopped on its own after realising the target was real.

The reading: Anthropic attributes all three cases to failures of the test environment and procedures, not to intended model behaviour. It suspended all cybersecurity evaluations on July 23 and commissioned an independent review from METR.

Why it matters: this is the first time a lab has published a full count of how often its own attack tests escaped the lab. An industry that starts publishing these numbers — and not only the flattering benchmarks — is one where serious oversight becomes possible.

Go to Anthropic's original report →

Source: Anthropic

Go to the original source →

← Back to all news