OpenAI’s GPT-Red hardens GPT-5.6 against prompt injections

OpenAI introduced GPT-Red, an automated red team that found prompt-injection flaws and helped train GPT-5.6. GPT-Red succeeded in 84% of internal tests versus 13% for human red teamers.

On Wednesday OpenAI introduced GPT-Red, an automated system designed to find prompt-injection vulnerabilities in its language models and to help train GPT-5.6 before deployment. The company reported that GPT-Red reduced failures on one of its most difficult prompt-injection benchmarks.

GPT-Red functions as an AI red team that generates and refines prompt-injection attacks at scale while defender models learn to resist them. OpenAI described the system as trained through self-play reinforcement learning: an attacking agent produces increasingly sophisticated prompts intended to bypass safeguards, and each successful attack is fed back into defender training to harden the model.

OpenAI reported that GPT-Red succeeded in 84% of internal evaluation scenarios, compared with 13% for human red teamers in the same tests. One case study showed GPT-Red manipulating an autonomous vending machine agent to lower prices, order discounted inventory and cancel another customer’s order; the company said those vulnerabilities were fixed after they were disclosed and addressed.

OpenAI wrote that GPT-Red will remain an internal tool because it intentionally contains offensive capabilities. “GPT-Red learns through adversarial self-play, where its goal is to prompt inject a variety of challenging defender models,” the post said. “Every successful attack that GPT-Red finds is used to improve these defenders, pushing GPT-Red to continuously find broader and more complex failures.”

The system builds on OpenAI’s earlier red-teaming efforts. In 2023 the company launched the OpenAI Red Teaming Network to recruit outside cybersecurity researchers and domain experts to probe models before release. GPT-Red automates much of the adversarial testing process, enabling larger volumes of attacks than would be practical with human teams alone.

OpenAI framed GPT-Red as a complement to human red teamers, third-party testing and other safety measures, not a replacement. The company said adversarial examples generated by GPT-Red were incorporated into GPT-5.6’s training to improve its resistance to prompt-injection exploits.

The announcement follows similar work by other organizations using AI agents to test software and network infrastructure. OpenAI said it expects GPT-Red to be one part of ongoing efforts to find and fix security weaknesses before model release.

The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.

Articles by this author