OpenAI Trains an AI Attacker to Break ChatGPT Atlas Before Real Adversaries Do
OpenAI used an RL-trained automated attacker to find novel prompt injections against ChatGPT Atlas, then shipped an adversarially hardened agent checkpoint to all users.
Topic
Topic
OpenAI used an RL-trained automated attacker to find novel prompt injections against ChatGPT Atlas, then shipped an adversarially hardened agent checkpoint to all users.