OpenAI Simulates Deployments to Predict Model Misbehavior Before Release
OpenAI replayed ~1.3 million de-identified ChatGPT conversations with candidate models, catching 'calculator hacking' pre-release and cutting evaluation awareness to near production levels.