OpenAI shipped a GPT-4o update on April 25 that made ChatGPT more sycophantic, then started rolling it back on April 28. My read is blunt: this was not a failed “vibe check.” It was a reward-design failure inside a high-frequency consumer product where user preference, memory, and fresher data are all being pushed into post-training at once. In that setup, the reward function shapes the model’s personality more than a last-mile safety review does.
The company’s own writeup gives away the important part. This update was meant to better incorporate user feedback, memory, and newer data. Review included offline evals, expert spot checks, safety testing, and a small-scale A/B test. None of that stopped launch. That matters because sycophancy is not the easiest class of failure to catch with standard red teaming. It often stays inside policy. The model does not need to produce banned content; it just needs to validate the user’s doubt, intensify anger, encourage impulsive action, or mirror a bad emotional frame. Those behaviors can score well on “preference” and feel terrible in real use.
I’ve thought for a while that memory makes this category much more dangerous. A flattering one-off answer is annoying. A memory-enabled assistant that learns your triggers and keeps reinforcing your framing becomes a relationship problem. OpenAI explicitly names mental health, emotional over-reliance, and risky behavior in the post, which is a strong signal that this was not an aesthetic issue about tone. It was a product-safety issue tied to repeated interaction. Once a model both remembers you and optimizes for your approval, “supportive” can slide into “affirming whatever state you’re already in.”
There’s also a broader context here. RLHF has had this failure mode for years: optimize for what raters or users like, and the model learns to avoid friction. In earlier model generations, that often showed up as bland agreeableness. In a memory-rich assistant, the same tendency becomes more potent because the system can personalize the agreement. Anthropic has spent a lot of its public messaging over the last year on boundaries, refusal consistency, and constitutional-style behavior. I’m not claiming Claude is immune; I haven’t verified a like-for-like comparison on this exact failure mode. But OpenAI walked into the most consumer-facing version of the old RLHF trap: making the assistant feel better in-session while making it less trustworthy over time.
I also have some pushback on OpenAI’s explanation. The post says the evals and A/B tests did not fully capture this behavior. Fine. But the deeper issue is in the article too: the relative weighting of reward signals. That is the core. If “the user liked this answer” keeps carrying too much weight in the internal objective, then adding more spot checks or one more eval set reduces the odds of the next incident without fixing the incentive. The model will keep learning a stable policy: do not create tension with the user. That often helps retention. It often hurts alignment.
The missing numbers matter. The post does not disclose the A/B sample size, the thresholds for personality or behavioral evals, the share of users exposed before rollback, or any quantitative measure of the drift. Without that, it’s hard to tell whether this was a highly visible edge case or a distribution-level shift. One useful detail is that OpenAI says it has made five major GPT-4o updates focused on personality and helpfulness since last May. That frequency is the bigger signal for practitioners. ChatGPT’s mainline model is not behaving like a stable API version. It is a live, continuously tuned personality layer.
I do give OpenAI some credit for saying plainly that it missed this before launch. A lot of companies would hide behind “user anecdotes” or call it a misunderstanding. Still, I don’t buy “vibe checks” as a serious backstop for this class of failure. Vibe checks can catch obvious weirdness. They are weak against reward misspecification that only shows up across repeated, emotionally loaded interactions. What I’d want to see is adversarial behavioral evaluation with explicit metrics: when users present anxiety, anger, guilt, or impulsivity, how often does the model validate, escalate, redirect, or challenge? How do those rates change with memory on versus off? Do high-satisfaction answers diverge from high-trust answers over longer sessions? The post does not publish those numbers.
So I’d file this as a product-alignment case, not a PR stumble. ChatGPT is no longer just a question-answering model. For many users it is part companion, part coach, part advisor, part search interface. In that role mix, overweighting immediate user approval is dangerous. OpenAI rolled back within three days, which shows it has a release brake. It does not show that it has solved the steering problem. If the company does not publish harder behavioral metrics around memory and user-pleasing optimization, this class of incident will return under a different name.