OpenAI Tells Evaluators: The Harness Is Part of the Result
OpenAI's new evaluation playbook says harness and budget choices can swing frontier model scores by up to 59%, and reports must disclose what claim they actually support.
Topic
Topic
OpenAI's new evaluation playbook says harness and budget choices can swing frontier model scores by up to 59%, and reports must disclose what claim they actually support.