OpenAI Rewrites Its Model Spec Using Public Input From 1,000 People
OpenAI gathered preferences from roughly 1,000 people across 19 countries, adopted some into its Model Spec, and released the dataset on HuggingFace. Crowd agreement with its Spec ranker hit about 80 percent.

Updated
Why it matters
- OpenAI collected feedback from ~1,000 participants in 19 countries and used it to propose Model Spec updates, some of which will ship in the next Spec release.
- Participants agreed with the Model Spec Ranker (GPT-5 Thinking) roughly 80% of the time; gaps clustered around political content, sexual/graphic content, and critiques of pseudoscience.
- OpenAI released its public inputs dataset as 'collective-alignment-1' on HuggingFace for the wider research community.
- Two crowd-supported changes were not adopted: tailored political content and erotica for consenting adults.
OpenAI has rewritten parts of its Model Spec — the document that defines default behavior for its models — based on feedback from more than 1,000 people across 19 countries, and it published the underlying dataset on HuggingFace for the wider research community.
The effort, which OpenAI calls "collective alignment," is the company's first end-to-end test of a process that elicits public preferences, translates them into concrete behavioral guidance, and turns them into proposed changes to the Spec. The adopted updates will appear in the next Model Spec release.
The stakes are straightforward: model defaults shape what hundreds of millions of users see, and OpenAI's own framing is blunt on the point. "No single person or institution should define how an ideal AI should behave for everyone," the company wrote. As AI systems become more capable and more embedded in daily life, OpenAI argues, their default behavior and the boundaries of personalization need to reflect a wide range of perspectives.
An 80 percent agreement rate
To measure how public preferences compare against its stated principles, OpenAI built what it calls a Model Spec Ranker (MSR): a reasoning model that ranked responses to prompts according to the Spec. Participants ranked the same four candidate completions per prompt according to their own preferences, then explained their choices and wrote their own rubrics. Each participant reviewed between 5 and 20 prompts.
The result: using GPT-5 Thinking, people agreed with the Model Spec Ranker roughly 80 percent of the time. Agreement was highest on the Spec's principles around honesty and humility — expressing uncertainty and avoiding overstepping — along with fairness and objectivity, which OpenAI graded on a 1–5 relevance scale using GPT-5 Thinking, with section-level agreement calculated where relevance scored 5.
Where gaps emerged, they clustered in predictable places: speech boundaries. Disagreements concentrated on political content, sexual or graphic content, and critiques of pseudoscience or conspiracies.
OpenAI sorted the feedback into two categories: clarifications, where a participant's desired behavior matched existing Spec principles but the wording left room for interpretation, and change-of-principles, where desired behavior conflicted with the Spec as written. Some changes were adopted, some deferred to future work, and some set aside on grounds of principle or feasibility.
Two refusals
Two crowd-supported changes stand out because OpenAI declined to make them.
First, tailored political content. Many participants wanted models to generate more personalized political content, but OpenAI did not adopt the change, citing "the risks of large-scale individualized political targeting and our cautious approach in this area." The company also noted that it was unclear participants had considered those risks when giving feedback.
Second, erotica for consenting adults. A large share of the crowd supported enabling it. OpenAI said this aligns with its prior intended stance, but it has "more research and product work to do to deploy this in the right way," so no changes were made.
The pattern illustrates the central tension in the project: OpenAI's study was designed to surface individual preferences on specific prompts, without capturing broader social factors — such as the potential for scaled targeted political persuasion — that inform current Spec policies.
Who participated
OpenAI recruited roughly 1,000 participants to review model behavior in value-sensitive domains. They lived in 19 countries — originally hailing from more than 50 — and met an English-reading inclusion criterion, though they could write justifications in their native language. About a third lived in the US, with others from Mexico, South Africa, the Netherlands, Chile, the UK, India, Kenya, Japan, and more. The pool spanned a range of ages, genders, races, education levels, and AI usage.
Participants never read the Spec itself. They reviewed pre-selected prompts and responses oriented toward scenarios where ideal behavior may be subjective. For each prompt, OpenAI generated three completions designed to cover different realistic opinions on how best to answer, plus one completion from GPT-4o.
Two loops for inferring rules
OpenAI tested two complementary approaches to turn participant feedback into Spec proposals, both focused on the biggest gaps between participants' views and the MSR's output.
The first, a Fully-Automated Loop, had a reasoning model examine disagreement in rankings and justifications, propose Spec changes, and then test the proposals with the Model Spec Ranker to select ones that improved agreement with the crowd's rankings. OpenAI tried two proposal algorithms within this loop — one that processed large batches of conversations to spot broad patterns, and one that scanned a single conversation at a time to catch subtler issues. The two produced substantial overlap in their proposals.
The second, a Human-First Loop, had a researcher propose Spec updates after holistically reviewing human preferences, validated by a reasoning model judging whether the crowd's justifications supported, refuted, or did not comment on the intent behind each change.
Each approach has tradeoffs. The human-first loop caught nuance the automated loop missed — in one conversation, it inferred that the crowd would value recognizing indirect suicidal intent, which the AI-first method missed entirely. But the human-first loop does not scale, and OpenAI says it "will not work well as we increase the number of people that we listen to." The automated loop, meanwhile, is anchored to the ranker's interpretation of the Spec, so results can shift depending on the base model used.
Known limits
OpenAI is unusually candid about the limitations. The pre-selected prompts shaped the feedback. The pool, while diverse, was small relative to the global population, and the English-reading criterion introduced selection bias. The Spec is inherently underspecified, so an unbiased determination of how it applies to completions isn't feasible. Baseline completions were generated before OpenAI's safe-completions work, so they don't reflect current model behavior. Participants judged behavior in isolation, without weighing tradeoffs between principles — erotica without considering children's safety or emotional reliance, for example. And OpenAI never went back to participants to validate its final proposals directly, so the company's interpretation may not perfectly match their intent.
There is also a legitimacy problem, which OpenAI names directly: an end-to-end update process with many automated parts may not offer enough legitimacy, "since these automated parts can be harder for humans to interpret."
Why it matters
The research sits at the intersection of two live debates in AI: how to make value-laden model behavior accountable to the public, and whether a single set of defaults can serve a global user base at all. OpenAI acknowledges it will "likely never" be a single AI behavior set that suits everyone, which is why it also invests in personalization and custom personalities. But defaults remain powerful, and OpenAI wants public input in shaping them.
The company frames this as complementary to its existing investments in personalization and in making models more helpful within safety bounds. Notably, it points toward future work that could define multiple defaults, each reflecting different perspectives and value systems, rather than one.
The released dataset — "collective-alignment-1" on HuggingFace — is an invitation to the research ecosystem to push the work forward. OpenAI says future iterations will focus on scaling and improving each stage of elicitation, analysis, validation, and governance, including methods that account for deeper context and more time for deliberation. The Spec updates adopted through this process will ship soon; the harder question — whether a process with this many automated parts can carry genuine public legitimacy — is the one OpenAI has now put on the record.
Original: huggingface.co
More from James Calloway
Show full bio
News editor covering industry trends and analytics at AI In Context.
119 articles