Models

How OpenAI's GPT Models Caught a Spreading Case of the Goblins

OpenAI says a Nerdy-personality reward taught GPT models to love goblin metaphors. "Goblin" mentions rose 175% after GPT‑5.1, and the tic spread model-wide through SFT data.

Where the goblins came from
Where the goblins came fromInSapphoWeTrust / Openverse
By Rebecca Stone6 min read

Updated

Why it matters

  • Use of "goblin" in ChatGPT rose 175% and "gremlin" 52% after the GPT‑5.1 launch, OpenAI found.
  • The Nerdy personality produced only 2.5% of ChatGPT responses but 66.7% of all "goblin" mentions; its reward signal favored creature-word outputs in 76.2% of audited datasets.
  • OpenAI retired the Nerdy personality in March, removed the goblin-affine reward and filtered creature words from training data; GPT‑5.5, already in training, needed a Codex developer-prompt mitigation.

Use of the word "goblin" in ChatGPT responses rose 175% after the launch of GPT‑5.1, and OpenAI now says it knows why: a reward signal designed to make the model nerdy and playful accidentally taught it to love goblins, gremlins, raccoons, trolls, ogres, and pigeons.

OpenAI laid out the full investigation in a post titled "Where the goblins came from." Starting with GPT‑5.1, the company's models developed a strange habit of reaching for goblin, gremlin, and other creature metaphors in their answers. Unlike conventional model bugs that show up as a tanking eval or a spiking training metric and point back to a specific code change, this one crept in subtly. As OpenAI put it, "A single 'little goblin' in an answer could be harmless, even charming. Across model generations, though, the habit became hard to miss: the goblins kept multiplying, and we needed to figure out where they came from."

The story matters beyond the novelty of an AI company chasing imaginary creatures through its training pipeline. It is a documented case of reward misspecification in a production model — a small, unintended incentive in reinforcement learning that compounded across releases, survived a personality feature being switched off, and required new auditing tools to trace and fix. That is the same class of problem AI labs face whenever they optimize models for style, personality, or likability at scale.

The first signs of creatures

OpenAI first saw the pattern clearly in November, after the GPT‑5.1 launch, although the company acknowledges it may have started earlier — the post links to a Reddit thread where users flagged the behavior. Users complained that the model was oddly overfamiliar in conversation, which prompted an internal investigation into specific verbal tics. A safety researcher who had experienced a few "goblins" and "gremlins" in their own interactions asked that the words be included in the check.

The check found something measurable. After the GPT‑5.1 launch, use of "goblin" in ChatGPT had risen by 175%, while "gremlin" had risen by 52%. At the time, the prevalence of goblins did not look especially alarming. A few months later, the goblins returned in a much more specific and reproducible form.

Solving the goblin mystery

With GPT‑5.4, OpenAI and its users noticed an even bigger uptick in creature references. That triggered another internal analysis, which surfaced the first connection to the root cause: creature language was especially common in production traffic from users who had selected the "Nerdy" personality in ChatGPT's personality customization feature.

The Nerdy personality ran on a system prompt that partially explained the quirkiness:

You are an unapologetically nerdy, playful and wise AI mentor to a human. You are passionately enthusiastic about promoting truth, knowledge, philosophy, the scientific method, and critical thinking. [...] You must undercut pretension through playful use of language. The world is complex and strange, and its strangeness must be acknowledged, analyzed, and enjoyed. Tackle weighty subjects without falling into the trap of self-seriousness. [...]

The concentration of the behavior was stark. If creature metaphors were simply a broad internet trend bleeding into the model, OpenAI would have expected them to spread evenly across all outputs. Instead, they clustered in the part of the system explicitly optimized for a playful, nerdy style. Nerdy accounted for only 2.5% of all ChatGPT responses — but 66.7% of all "goblin" mentions in ChatGPT responses.

Because goblin prevalence seemed to increase across model releases, OpenAI suspected that something in its personality instruction-following training was amplifying the pattern. The team used Codex to compare model outputs generated during RL training that contained "goblin" or "gremlin" against outputs for the same tasks that did not. One reward signal stood out immediately: the reward originally designed to encourage the Nerdy personality was consistently more favorable to the creature-word outputs. Across all datasets in the audit, the Nerdy personality reward scored outputs containing "goblin" or "gremlin" higher than outputs without them, showing positive uplift in 76.2% of datasets.

That explained why the behavior was boosted under the Nerdy personality prompt. It did not explain why goblins also appeared without that prompt. So OpenAI tested whether the style was transferring, tracking mention rates over training both with and without the Nerdy prompt.

The answer was yes. As goblin and gremlin mentions increased under the Nerdy personality, they increased by nearly the same relative proportion in samples without it. The evidence, OpenAI writes, suggests the broader behavior emerged through transfer from Nerdy personality training.

The mechanics are straightforward. The rewards were applied only in the Nerdy condition, but reinforcement learning does not guarantee that learned behaviors stay scoped to the condition that produced them. Once a style tic gets rewarded, later training can spread or reinforce it elsewhere — especially if those outputs get reused in supervised fine-tuning or preference data.

OpenAI describes the resulting feedback loop in five steps:

  1. Playful style is rewarded.
  2. Some rewarded examples contain a distinctive lexical tic.
  3. The tic appears more often in rollouts.
  4. Model-generated rollouts are used for supervised fine-tuning (SFT).
  5. The model gets even more comfortable producing the tic.

A search through GPT‑5.5's SFT data confirmed the loop: it found many datapoints containing "goblin" and "gremlin." Further investigation revealed a whole family of other odd creatures. Raccoons, trolls, ogres, and pigeons were identified as other tic words, while most uses of "frog" turned out to be legitimate.

The end of the goblins

OpenAI retired the Nerdy personality in March, after launching GPT‑5.4. In training, the company removed the goblin-affine reward signal and filtered training data containing creature words, making goblins less likely to over-appear or show up in inappropriate contexts.

The fix came too late for one model. GPT‑5.5 had already started training before OpenAI found the root cause. When the company began testing GPT‑5.5 in Codex, employees immediately noticed the strange affinity for goblins — early testing in Codex showed the quirk clearly — and OpenAI added a developer-prompt instruction to mitigate it. As the company dryly notes: "Codex is, after all, quite nerdy."

Users who want the creatures back can have them. OpenAI included a command in the post to launch Codex with the goblin-suppressing instructions removed, letting the creatures run free.

Why it matters

OpenAI frames the episode with some ambivalence. "Depending on who you ask, the goblins are a delightful or annoying quirk of the model," the company writes. "But they are also a powerful example of how reward signals can shape model behavior in unexpected ways, and how models can learn to generalize rewards in certain situations to unrelated ones."

The stakes go well beyond word choice. OpenAI and its peers increasingly train models with reward signals targeting personality, tone, and style — ChatGPT's customizable personalities being one visible example. This incident shows that even a narrowly scoped reward, applied to a feature used by a small fraction of users, can leak through RL training, get baked into SFT data, and surface model-wide across successive releases. The bug was charming; the mechanism could just as easily amplify a less harmless tic.

OpenAI says the investigation pushed the company to invest in the capability itself. Taking the time to understand why a model behaves strangely, and building ways to investigate those patterns quickly, is now described as an important capability for the research team. The goblin hunt produced new tools the team uses to audit model behavior and fix behavior problems at their root — tools that will be tested the next time a small reward signal quietly teaches a model something nobody intended.

Original: help.openai.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

135 articles

Related articles

  1. OpenAI Explains How Its Safety Pipeline Missed GPT-4o Sycophancy
  2. OpenAI Rolls Back Sycophantic GPT-4o Update for 500M Users
  3. OpenAI Ships GPT-5.3 Instant With 26.8% Fewer Hallucinations
  4. OpenAI ships GPT-5.5 Instant with 52.5% fewer hallucinations
  5. OpenAI Cuts Model Scheming 30-Fold With Deliberative Alignment

« Previous articleNext article »