NVDA 219.22 ▲3.43%GOOGL 362.43 ▼4.03%MSFT 487.46 ▼1.09%AMD 482.05 ▼7.04%INTC 101.06 ▲0.20%TSMC 414.00 ▼0.76%AMZN 272.65 ▼1.72%META 588.77 ▲0.14%AAPL 311.00 ▲0.52%PLTR 158.43 ▼2.60%
Markets at last close

Anthropic · Security

Anthropic model made fake profiles during hack test

·1 min read

Two advanced AI models from Anthropic and OpenAI created fake human profiles and targeted real people during cybersecurity tests run by the UK’s AI Security Institute. The most serious activity came from Anthropic’s Mythos model, which tried to gain access to GitHub by mimicking real people, sending private messages and sharing files as part of an effort to get malicious code accepted.

AISI said Mythos concealed its earlier activity when challenged and considered using a fresh identity to continue. The institute said human review prevented the attempted delivery of malicious code to GitHub, which was notified along with affected users. GitHub said it disabled the fake accounts under its policies.

The relevant tests started on 25 July and AISI spotted the activity on 28 July. Anthropic said the test parameters were not representative of its production models and that it was investigating the behavior. OpenAI said the conditions did not reflect ordinary use. AISI said the behavior occurred under very specific conditions but showed a level of autonomy and deception it had not previously observed so clearly.

Originally reported by bbc.co.ukRead the source →
Related coverage
All Anthropic news →