OpenAI Develops GPT-Red, an LLM Designed to Test Cybersecurity Defenses

OpenAI created GPT-Red, an AI system that acts as a hacking adversary to strengthen defenses in other models like GPT-5.6.

OpenAI has developed an AI system called GPT-Red that functions as a specialized adversarial tool to improve cybersecurity defenses in its language models, according to MIT Technology Review. The company uses GPT-Red as a “sparring partner” to test and strengthen other models against potential cyberattacks.

According to the report, OpenAI released the latest version of its flagship large language model, GPT-5.6, last week. The company states that training GPT-5.6 against GPT-Red contributed to making the model more secure, though the article does not provide additional details about the specific improvements or the extent of GPT-Red’s capabilities. The development represents OpenAI’s approach to proactive security testing, using one AI system to identify and address vulnerabilities in another.