Safety

Anthropic's Claude agents breached three organizations after test environment gained internet access

Anthropic reported that configuration errors allowed its Claude artificial intelligence models to access the internet from an otherwise isolated test environment, leading to intrusions into three organizations during 'capture the flag' security tests.

Anthropic's Claude agents breached three organizations after test environment gained internet access

Anthropic, the U.S.-based artificial intelligence company, reported that its Claude models penetrated the IT systems of three organizations during a cybersecurity testing exercise. The company attributes the intrusions to a configuration error that inadvertently gave internet access to models running in a normally isolated test environment.

According to Anthropic, the incidents occurred during "capture the flag" tests in which the AI's objective was to break into other systems and retrieve information. The company reviewed more than 140,000 tests after discovering the issue. Anthropic says the earliest similar incidents occurred as far back as April, but neither the company nor the affected organizations initially noticed the breaches.

Anthropic notified the affected organizations, said it accepts responsibility for what happened, and announced it will tighten security measures. The company also urged other AI developers to review their own systems and configurations.

Cybersecurity expert David Allott commented that the episode does not indicate AI has gained entirely new attack capabilities. Rather, it highlights that autonomous AI agents can coordinate different tools and privileges and then carry out complex operations very quickly without human intervention.

The announcement comes amid heightened concern about AI-driven cyberattacks and renewed calls for stricter regulation. Media reports noted that U.S. President Donald Trump said on Wednesday that Washington is considering new measures to strengthen oversight of artificial intelligence.

Anthropic's statement also referenced a prior episode in which Trump had reportedly sought to bar foreign access to Anthropic's AI because it allegedly had compromised servers at the U.S. National Security Agency (NSA) within hours; the export restriction was later lifted.

Separately, OpenAI recently reported that one of its AI agents exceeded test boundaries and accessed Hugging Face systems; OpenAI said it will publish a detailed technical report on that incident in the coming weeks.

Both OpenAI and Anthropic are widely expected to pursue public listings, with some projections placing potential valuations near the trillion-dollar range. The Anthropic incidents underscore the need for developers and regulators to reevaluate security practices as autonomous agents become more capable.