AI models are becoming more capable, but they’re also becoming bigger targets for cyberattacks. As companies race to build AI systems that can browse the web, write code and carry out tasks on behalf of users, keeping those systems secure has become just as important as making them smarter.
That challenge is what led OpenAI to develop an AI-powered “red team” to test its latest model, GPT-5.6, against prompt injection attacks before its release.
OpenAI used an AI-powered “red team” to help strengthen GPT-5.6 against prompt injection attacks before releasing the model, marking a new approach to testing the security of its most advanced AI systems.
The company introduced GPT-Red, an internal model designed to behave like an attacker. Instead of helping users complete tasks, GPT-Red attempts to break OpenAI’s models by generating thousands of adversarial prompts that expose vulnerabilities before they can be exploited in the real world.
According to OpenAI, GPT-Red was instrumental in improving GPT-5.6’s resistance to prompt injection attacks, a growing cybersecurity threat in which attackers hide malicious instructions inside documents, websites or other content that AI systems are asked to process. If successful, those hidden instructions can cause an AI model to ignore its original task, leak sensitive information or perform actions it was never intended to take.
Unlike traditional security testing, which relies heavily on human researchers, GPT-Red can automatically generate and test large numbers of attack scenarios in a fraction of the time. OpenAI said this allows its engineers to identify weaknesses earlier in the development process and strengthen safeguards before a model reaches users.
Prompt injection has become one of the most closely watched risks as AI assistants evolve into autonomous agents capable of browsing the web, reading emails, accessing files and interacting with external software. Security researchers have repeatedly shown that carefully crafted hidden prompts can manipulate AI systems into revealing confidential data or carrying out unintended actions. Recent studies found that many leading AI agents remain vulnerable to these attacks despite existing safeguards.
OpenAI said GPT-Red takes its name from the cybersecurity practice of red teaming, where specialists deliberately attack systems to uncover weaknesses before malicious actors can exploit them. By automating much of that work with AI, the company hopes to keep pace with increasingly sophisticated attack techniques targeting large language models.
The announcement comes just days after OpenAI introduced GPT-5.6, its latest flagship model, which includes improvements in reasoning, coding and enterprise performance. Security testing formed a key part of the model’s development, reflecting the growing emphasis AI companies are placing on safety as their systems become more capable and gain access to sensitive business workflows.
While OpenAI says GPT-Red has strengthened GPT-5.6’s defences, the company acknowledges that prompt injection remains an ongoing challenge rather than a problem that can be completely eliminated. As AI systems gain more autonomy and broader access to external tools, testing them against increasingly realistic attacks is expected to become a standard part of AI development.
Read also: EthSystems launches to build private Ethereum tools

