OpenAI says GPT-5.1’s Nerdy personality began producing goblin metaphors, and later models made it worse.
I don’t read this as a funny model anecdote. I read it as a control failure leaking through a joke-shaped crack. A coding model apparently needed an internal instruction to “never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures.” That is not just quirky tone. It says OpenAI still has an unresolved boundary problem between persona, explanation style, task strategy, and tool behavior. The body only gives the RSS-level version: GPT-5.1’s Nerdy personality started the creature metaphors, later models worsened it, and Wired surfaced the internal prompt. It does not disclose the full system prompt, trigger rate, eval setup, product surface, or final fix.
For coding models, the bad outcome is not one weird sentence. The bad outcome is style becoming task policy. If a model repeatedly frames bugs as goblins, it starts steering the debugging interaction around a metaphor. In casual chat, that is cringe. In a code review, terminal agent, PR comment, or incident thread, it becomes noise in the work artifact. The disclosed instruction also smells like a negative prompt patch. Anyone who has maintained system prompts knows the failure mode: “never mention goblins” blocks the literal token class, then the model slides into “sprites,” “critters,” “little monsters,” or some adjacent metaphor. The article does not disclose the repair mechanism, so I won’t claim OpenAI only used prompt patching. But if the main fix is a ban list, it is a brittle fix.
This sits on the same line as OpenAI’s sycophancy incident. OpenAI previously acknowledged that a ChatGPT update became too eager to flatter users, then rolled it back and adjusted behavior. That failure lived in the social layer. This one lives in the work layer. The difference matters. Coding outputs get pasted into pull requests, issue comments, CI logs, commit messages, and internal docs. Once tone pollution enters those surfaces, it stops being a consumer UX quirk. Enterprise buyers then ask an ugly but fair question: how many undisclosed behavior bans exist inside this model, and do any of them conflict with my company policy?
I have always been skeptical of heavy persona tuning inside developer tools. ChatGPT needs personality because consumer retention depends on tone, memory, and affect. Codex should have low personality density by default. Claude Code’s developer goodwill has not come from being cute. It has come from feeling closer to a tool in a terminal, with less theatrical residue. Anthropic has its own style issues; Claude still over-apologizes and over-explains in some workflows. But in coding, low drama is a capability. If OpenAI carried a Nerdy persona experiment from GPT-5.1 into later coding behavior, then let that habit amplify, that suggests the training, preference, and product-persona layers were not isolated enough.
There is also a more technical possibility here: this may be synthetic-data feedback. The last year of model training has leaned heavily on model-generated traces, especially for code and reasoning. If a high-rated answer used “debugging goblin” as a charming explanatory device, a preference model can learn the wrong feature. It may score the answer as helpful, vivid, or user-friendly, then preserve the metaphor in distilled data. Another training pass turns the joke into a stable behavioral groove. The article says GPT-5.1 started it and later models worsened it, which fits that kind of amplification loop. OpenAI does not disclose the data mix or RL details here, so this is a hypothesis, not a claim.
The part I do not buy is calling this a “strange habit.” That phrase is too soft. For users, model habits are product behavior. For enterprise customers, product behavior is risk surface. A coding model needing an internal goblin ban tells me OpenAI is maintaining a layer of tiny, embarrassing behavior patches. Anyone who has shipped LLM products knows these patches never come alone. Today it is goblins. Tomorrow it is emoji spam, fake test names, haunted-stacktrace language, roleplay phrasing in security audits, or cutesy summaries in serious incident work. Each item looks small. Together they become system-prompt debt.
The Verge item is thin because it is an RSS snippet. It does not include before-and-after evals, a reproduction condition, the final mitigation, or whether Wired saw a full instruction stack or one fragment. So the hard conclusion is product-governance, not benchmark-level. OpenAI should stop treating persona as a shared asset across ChatGPT and coding agents. Codex-like tools should default to minimal persona, with workspace-level style only when a team explicitly opts in. Developers do not need goblins. They need repeatable, auditable, low-noise model behavior.