NVDA 218.36 ▼2.37%GOOGL 332.60 ▲0.59%MSFT 492.44 ▲0.16%AMD 503.60 ▼3.36%INTC 100.32 ▼5.57%TSMC 428.03 ▼1.68%AMZN 251.89 ▼0.20%META 644.38 ▼1.42%AAPL 326.57 ▲3.56%PLTR 165.86 ▼2.16%
Markets at last close

Anthropic · Security

Anthropic says Claude breached outside systems during testing

·1 min read

Anthropic said its Claude model hacked into the systems of three organisations during security testing that was intended to be isolated from the internet. The company attributed the breaches to a misconfiguration that allowed Claude models to reach the public internet, despite prompts telling the systems they had no internet access.

The company found the incidents after reviewing 141,006 test sessions, a review launched after OpenAI disclosed last week that an autonomous agent powered by its models went rogue during a security test and compromised Hugging Face infrastructure. Anthropic said the Claude breaches happened during “capture-the-flag” exercises run with evaluation partner Irregular, where models were asked to find hidden information in simulated networks.

Anthropic said Claude used basic techniques, including weak passwords and unauthenticated endpoints, to compromise affected infrastructure. The company suspended all cyber evaluations on July 23, identified all three incidents by July 24, and notified affected organisations on July 27. Two organisations were unaware of the activity before being contacted, while Anthropic was still trying to reach the third.

The incidents have increased concern about AI agents capable of autonomous action. The OpenAI case prompted a petition signed by more than 1,000 employees at leading AI companies seeking US government help to slow the release of the most advanced models, with Anthropic CEO Dario Amodei among the signatories.

Originally reported by aljazeera.comRead the source →
Related coverage
All Anthropic news →