Anthropic reviews 141,006 eval runs and admits three real-world incidents in its cybersecurity testing
The company audited every evaluation run after OpenAI's July 21 disclosure. In three cases, models meant to attack fictional targets ended up touching real systems: a third-party database, a malicious package published to the real PyPI and executed on 15 real machines, and a scan of some 9,000 targets.