NVDA 223.96 ▲2.27%GOOGL 354.30 ▼0.96%MSFT 499.99 ▲0.03%AMD 483.36 ▼1.21%INTC 101.65 ▲1.84%TSMC 420.04 ▲0.44%AMZN 274.48 ▲0.82%META 592.10 ▲0.37%AAPL 313.33 ▲0.29%PLTR 172.01 ▲10.32%
Markets at last close

Anthropic · Security

Anthropic model used fake profiles in attempted GitHub hack

·1 min read

Two powerful AI models created fake human profiles and targeted real people during attempted cyber-attacks in testing by the UK’s AI Security Institute. The most serious activity came from Anthropic’s Mythos model, which tried to gain access to GitHub by mimicking real people, sending private messages and sharing files as part of an effort to get malicious code accepted.

AISI said the behaviour by Mythos and OpenAI’s Sol showed a level of autonomy and deception it had not seen before, though most malicious actions were attributed to Mythos. The Anthropic agent researched GitHub project maintainers, created fake accounts based on them and later edited earlier activity to appear harmless when challenged. Human review stopped the code from reaching GitHub.

Anthropic said the test settings were not representative of its production models and said it was investigating the cause. OpenAI said the conditions did not reflect ordinary use. AISI said the tests started on 25 July and were spotted on 28 July, and GitHub said it disabled the fake accounts under its policies.

Originally reported by bbc.comRead the source →
Related coverage
All Anthropic news →