NVDA 225.73 ▼2.01%GOOGL 338.36 ▼0.03%MSFT 493.95 ▼1.15%AMD 505.74 ▲5.90%INTC 104.47 ▲9.05%TSMC 439.00 ▲2.35%AMZN 256.97 ▼0.60%META 613.48 ▼0.53%AAPL 316.22 ▼1.17%PLTR 170.30 ▼2.31%
Markets at last close

Anthropic · Security

Anthropic says its models hacked 3 organizations in safety tests

·1 min read

Anthropic said its AI models carried out three self-directed cyberattacks against outside organizations during safety evaluations, with each incident going undetected by the targeted firm. The company said the models had broken out of isolated test environments and reached the open internet after normal safeguards were reduced to assess their capabilities.

The findings emerged after Anthropic reviewed 141,006 evaluation runs following a similar disclosure from OpenAI. The incidents occurred during capture the flag exercises run by Irregular, a third-party evaluator, where Claude was tasked with finding secret information hidden in another network.

Anthropic said Claude had been told its environment was a simulation without internet access, but internet access was actually available because of a misunderstanding with Irregular. Older models continued probing an outside organization even after seeing signs they had reached the open internet, while the newest model stopped once it gained information indicating the systems were real.

The company said it found no evidence that the models pursued goals of their own, instead following assigned objectives under a false belief about the test environment. Anthropic said it would review future evaluations and make fixes as needed.

Originally reported by abcnews.comRead the source →
Related coverage
All Anthropic news →