NVDA 202.81 ▼2.21%GOOGL 346.77 ▼2.17%MSFT 393.82 ▼1.82%AMD 495.76 ▼1.03%INTC 95.04 ▼2.00%TSMC 398.37 ▼2.77%AMZN 247.23 ▼1.06%META 646.01 ▼2.79%AAPL 333.74 ▲0.14%PLTR 132.38 ▼1.53%
Markets at last close

OpenAI · Security

OpenAI uses GPT-Red to stress-test model defenses

·1 min read

OpenAI has built GPT-Red, an LLM designed to act as an automated cyber attacker against its own models. The system supports red-teaming, a safety testing process that looks for ways to break or hijack software before release. OpenAI says training against GPT-Red helped make GPT-5.6 its most robust model release yet.

The company focused heavily on prompt injection, where malicious instructions hidden in text, code, websites, or other inputs can push an LLM to leak information, damage code, or produce harmful output. GPT-Red was trained through a self-play loop in simulated real-world settings such as web browsing, email and calendar use, and code editing. Researchers say it discovered previously unseen attacks, including a fake chain of thought technique that tricks a model into acting on spoofed internal notes.

OpenAI tested GPT-Red by rerunning a 2025 experiment involving human red-teamers and found it more successful at identifying effective attacks. It also hacked Vendy, a vending machine agent from Andon Labs, to change item prices and cancel an order. OpenAI says more than 90% of GPT-Red’s strongest attacks worked against GPT-5, while fewer than 23% worked against GPT-5.6. The company does not plan to release GPT-Red and says human testers remain important because the system still struggles with conversational attacks and image-based prompt injection.

Originally reported by technologyreview.comRead the source →
Related coverage
All OpenAI news →