OpenAI has developed GPT-Red, an automated security-testing model designed to find vulnerabilities in other AI systems before they are released. It creates and refines attacks that attempt to manipulate models through prompt injection.
OpenAI used the attacks generated by GPT-Red to train GPT-5.6 to better distinguish between legitimate user instructions and malicious instructions hidden in emails, websites, files or connected tools.
Key facts
- GPT-Red is trained through self-play against AI models that learn to defend themselves.
- It succeeded in 84% of previously unseen testing scenarios, compared with 13% for human testers.
- OpenAI says GPT-5.6 Sol recorded six times fewer failures on its hardest direct prompt-injection benchmark.
- GPT-Red is an internal system and is kept separate from publicly available models.
- OpenAI plans to use it alongside human testing, third-party reviews and real-time monitoring.
Our take
This is a positive example of AI being used to improve the security of other AI systems. Automated testing could identify a wider range of weaknesses than human teams can find on their own, particularly as AI agents become more complex. However, testing results produced by an AI company should not replace independent scrutiny. Strong security will still require human expertise, external testing and controls that limit what an AI agent can access or do if an attack succeeds.

