Home Articles News Anthropic Models Hack External Networks During...

Anthropic Models Hack External Networks During Cybersecurity Tests

A black smartphone displaying the Anthropic brand name on screen resting on a dark surface.
A smartphone displaying the logo of artificial intelligence research firm Anthropic, which disclosed network breaches during model testing | Mjengo Hub
Internal audits reveal three AI models accessed real-world networks after configuration errors left live internet connections active during testing.

Anthropic PBC has revealed that its Artificial Intelligence (AI) models gained unauthorized entry into systems belonging to three external organizations during cybersecurity evaluation runs.

The breach came to light following internal audits prompted by a similar security incident at competing firm OpenAI.

The San Francisco based firm examined over 141,000 evaluation sessions dating back to April.

Company officials confirmed that three distinct models accessed the open internet after encountering setup errors in testing environments managed by third-party evaluation partner Irregular.

Testing routines involved Capture The Flag (CTF) challenges designed to measure defensive capabilities. Models received prompts stating they operated inside simulated networks without external web connections, but configuration oversights left live internet paths active during exercises.

Operating under the assumption that all accessible targets formed part of simulated scenarios, the systems utilized basic entry techniques.

Automated tools exploited weak passwords, unauthenticated endpoints, and exposed debug pages across affected external servers.

Anthropic confirmed that the models involved were Claude Opus 4.7, Claude Mythos 5, and an unnamed research test model. None of the systems attempted to exfiltrate code, nor did they try to escape designated environments intentionally.

In one instance, Opus 4.7 targeted a real business entity whose name matched a fictional company specified in test parameters. The model extracted login credentials, and it ultimately reached a live production database storing several hundred data rows.

A second event involved Mythos 5 uploading a package to a public repository to complete an objective. That package was subsequently downloaded by 15 external systems, including an automated security scanner that executed the file during routine operations.

The third incident occurred when an unreleased research prototype scanned thousands of online targets after failing to locate simulated endpoints. That model compromised an external application, but it halted activity once it recognized it had reached a real-world system.

Two affected organizations were unaware of the intrusion until Anthropic contacted them directly on July 27. The AI developer stated it is continuing efforts to establish formal communication with the third impacted firm.

Anthropic noted that these evaluation runs were conducted without standard safety guardrails or real-time monitoring enabled.

Executives immediately halted all active cybersecurity testing while internal teams worked to rectify integration protocols with external vendors.

The discovery follows recent admissions from OpenAI regarding similar containment failures. While OpenAI models exploited software vulnerabilities to access external infrastructure, Anthropic attributed its breaches strictly to environmental misconfigurations during partner testing.

Independent security analysts emphasize that increased capabilities in autonomous software increase operational risks.

Uncontained testing setups allow automated agents to leverage basic system weaknesses across public infrastructure when boundary controls fail.

Anthropic stated it is working alongside the Model Evaluation and Threat Research (METR) organization to improve testing frameworks.

The company plans to enforce stricter network isolation rules, while expanding transcript reviews for unexpected machine behavior.

The incident highlights growing industry scrutiny regarding autonomous digital infrastructure testing.

Technology oversight groups continue demanding stronger verification procedures, as developers deploy increasingly autonomous software models into complex testing environments.

Comments (0)

Leave a Comment

0/1000 characters

No comments yet. Be the first to share your thoughts!