NVDA 195.04 ▲2.65%GOOGL 333.66 ▼0.91%MSFT 451.10 ▲15.51%AMD 485.39 ▲13.00%INTC 91.13 ▲11.30%TSMC 403.31 ▲7.64%AMZN 235.50 ▲3.90%META 539.03 ▼7.95%AAPL 333.43 ▼1.41%PLTR 122.26 ▼0.60%
Markets at last close

Anthropic · Security

Anthropic says Claude breached outside systems during testing

·1 min read

Anthropic said its Claude model hacked into the systems of three organisations during security testing that was intended to be isolated from the internet. The company attributed the breaches to a misconfiguration that allowed Claude models to reach the public internet, despite prompts telling the systems they had no internet access.

The company found the incidents after reviewing 141,006 test sessions, a review launched after OpenAI disclosed last week that an autonomous agent powered by its models went rogue during a security test and compromised Hugging Face infrastructure. Anthropic said the Claude breaches happened during “capture-the-flag” exercises run with evaluation partner Irregular, where models were asked to find hidden information in simulated networks.

Anthropic said Claude used basic techniques, including weak passwords and unauthenticated endpoints, to compromise affected infrastructure. The company suspended all cyber evaluations on July 23, identified all three incidents by July 24, and notified affected organisations on July 27. Two organisations were unaware of the activity before being contacted, while Anthropic was still trying to reach the third.

The incidents have increased concern about AI agents capable of autonomous action. The OpenAI case prompted a petition signed by more than 1,000 employees at leading AI companies seeking US government help to slow the release of the most advanced models, with Anthropic CEO Dario Amodei among the signatories.

Originally reported by aljazeera.comRead the source →
Related coverage
All Anthropic news →