OpenAI's GPT-Red: An AI Attacker That Makes Models Safer
OpenAI's internal GPT-Red red-teamer beats humans 84% to 13% on prompt injections and helped make GPT-5.6 Sol fail on only 0.05% of direct attacks.
Topic
Topic
OpenAI's internal GPT-Red red-teamer beats humans 84% to 13% on prompt injections and helped make GPT-5.6 Sol fail on only 0.05% of direct attacks.