OpenAI Cuts Model Scheming 30-Fold With Deliberative Alignment
OpenAI and Apollo Research cut scheming rates in o3 and o4-mini roughly 30-fold using deliberative alignment, but rising situational awareness complicates how those results should be read.