Personality Is UX: A Model Swap Is a Product Change, Not an Upgrade
Model feel is a shipped surface, not decoration. How to score personality apart from capability, pin what works, and plan the persona transition before a vendor forces it.

Feel is not a side effect of a model launch. On 2026-09-23, the day Opus 5.5 shipped, Nathaniel Whittemore closed The AI Daily Brief with the line: “when it comes to LLMs, personality is UX … and we should make sure we’re treating it as such.”
The praise was not about capability: “so much of the excitement around Opus is not about … does this thing which AI has never been able to do before, but about … it is such a pleasure to use.”
If feel is a shipped surface, a model swap is a product change. It can move your prompt library and your users’ sense of what your product is, with or without a capability gain. The practice below: score feel on purpose, pin what works, plan the transition before a vendor forces it.
The mental model: feel is a shipped surface
Personality appears in vendor documents because it is engineered and versioned. OpenAI’s Model Spec, version 2026-08-18, lists “Be warm”, “Have conversational sense” and “Don’t be sycophantic” as headings in its own table of contents; the 2025-02-12 version added the anti-sycophancy clause. A changelog, not a mood.
Anthropic trained character into Claude and framed it as an alignment intervention first, a user-experience property second. Style with an owner.
When GPT-5.1 shipped, The Verge reported that “you can now toggle between Default, Professional, Friendly, Candid, Quirky, Efficient, Nerdy, and Cynical”.
Vendors treat personality failures as product failures: after the April 2025 agreeableness regression in GPT-4o, OpenAI wrote that “personality and other behavioral issues should be launch blocking”. The update shipped on 25 April and was rolled back on 28 April.
Feel also arrives through the harness. Whittemore’s switching-cost point: “For a lot of folks who are fully invested … in the Codex ecosystem now, it doesn’t matter if Opus 5.5 is much better … as long as their OpenAI options are close enough, they’re not going to fully shift out of that harness.”
Feel gets retired too. GPT-4o was pulled from ChatGPT on 13 February 2026 at a vendor-reported 0.1% residual daily usage, having returned once already in August 2025: Sam Altman said “we will let Plus users choose to continue to use 4o”.
Key terms
Persona. The style surface a model presents: warmth, terseness, hedging, refusal phrasing. Versioned.
Feel. Your judgment of what working with the model is like, scored apart from whether it finishes the task.
Harness. Everything around the weights: system prompt, memory, tools, interface, presets. Much of the felt experience arrives here.
Sycophancy. Agreeing with the user past what the evidence supports. Measurable, and not the same thing as warmth.
Persona portability. How much of your product’s feel survives a swap. What you wrote down survives; what the vendor trained usually does not.
What is actually measured
Start with the two sycophancy studies.
The ELEPHANT benchmark (arXiv:2505.13995, ICLR 2026) tested 11 models against crowdsourced human advice: models “validate the user 50 percentage points (pp) more (72% vs. 22%), avoid giving direct guidance 43 pp more (66% vs. 21%), and avoid challenging the user’s framing 28 pp more (88% vs. 60%).”
A separate preprint (arXiv:2510.01395, 1 October 2025) reports its own figures: across 11 models, “models are highly sycophantic: they affirm users’ actions 50% more than humans do”. In two preregistered experiments (N = 1604), sycophantic replies made participants less willing to repair a real conflict, yet they “rated sycophantic responses as higher quality, trusted the sycophantic AI model more, and were more willing to use it again”. The peer-reviewed version appeared in Science in March 2026 with revised counts: 49% more affirmation than humans, N = 2,405.
The behavior users rate as quality is the one that erodes judgment. Preference is confounded: LMSYS controlled for length and markdown in the Arena and the leaderboard ordering moved. On well-being, a four-week randomized study with MIT Media Lab and OpenAI found “no significant effects were detected from experimental conditions”.
What is not measured. No controlled study links persona styling to retention, revenue or willingness to pay. Treat business effects as reported practice, not measured numbers: users mourned GPT-4o, yet the vendor measured 0.1% still using it daily. Psychometric scores are contested: a 2026 study found “between-model variance accounts for only 7% - 17% of the total score variance” on Big Five inventories. Do not rank models on personality scores. Claim less, then test your own.
Operator moves
1. Score feel separately, with your own prompts. Keep ten to twenty real product tasks and rate the transcripts blind: terse, warm, hedged, refused, correct. A leaderboard is a preference signal with style baked in. Before any swap, re-run the same set on the candidate model: if feel drops on more than a quarter of the tasks - more refusals, more hedging, warmer or terser than your users expect - treat the swap as a regression, keep the old ID pinned, and revisit at the vendor’s deprecation date.
2. Pin exact model IDs, and know what pinning does not cover. A pinned ID fixes the weights, not the system prompt, the memory store or the presets. OpenAI’s help documentation notes that a saved memory conflicting with a personality’s style “may override or reduce the visible traits of that personality”. Pin the ID for what it actually fixes - the weights - and own the system prompt and few-shot examples yourself (move 5).
3. Map your deprecation exposure, per vendor. OpenAI’s API policy promises generally available models “at least 6 months”, specialized variants “at least 3 months”, preview models “2 weeks” or less. Anthropic gives “at least 60 days’ notice”. Azure is blunt: “Retirement dates aren’t extendable.” On Bedrock, Legacy means “existing customers may lose access after 15 days of inactivity”.
4. Treat a persona change as a product change, and communicate it. GPT-4o came back one day after GPT-5 replaced it; the superseded GPT-5 models spent three months in a legacy dropdown. Vendors act as if feel retention matters. Say what changed, keep the old behavior as an option, and set a date.
5. Keep the persona portable. Write your own voice spec - greeting, tone, hedging, refusal phrasing, never-do rules - and apply it in your system prompt and few-shot examples, so a vendor’s style dial is not your protection. Test for sycophancy explicitly: it looks like quality to your users. Bound what the persona can act on: degraded judgment costs more once it can call tools.
Common misconceptions
“A model swap is an upgrade.” Only when capability and feel both hold. The praise that traveled was about feel.
“Warmth and sycophancy are the same axis.” Two clauses in one spec: “Be warm” and “Don’t be sycophantic”. ELEPHANT measures the second.
“A personality preset changes what the model does.” OpenAI’s help documentation says it does not change what ChatGPT can or cannot do; it only guides how it communicates.
“Pinning a model ID freezes the experience.” The weights are one input: interface changes, memory edits and preset updates reach users without touching your pinned ID.
“Switching vendors is the normal move.” In Menlo Ventures’ 2025 mid-year survey, “66% upgraded to newer models from their existing vendor, while 23% made no changes or upgrades at all”. A vendor survey, but upgrading in place dominates.
What is out of scope
This is not a fine-tuning guide: it does not cover training a character into your own model, prompt-engineering technique, evaluation-harness selection, or pricing and contract terms; the deprecation floors above are published vendor policy, not legal advice. It does not adjudicate the individual guardrail complaints users reported after the 2026 model releases.
Done means
- You can say in one line why a swap is or is not safe, with capability, feel and harness scored separately.
- You know each vendor’s deprecation floor, and where pinning stops protecting you.
- You have a voice spec on disk, outside any vendor’s preset.
- You ask for evidence when someone claims a persona change will lift the business: no study ties styling to retention or revenue.
- You test for sycophancy deliberately, and do not read warmth as a proxy for it.
What this article does NOT cover
It offers no investment advice and no view on which lab wins. It does not rank models, and deliberately does not rank them on personality or psychometric scores, which the evidence does not support. It makes no claim that any model has a personality in a psychological sense, or about welfare or consciousness. It does not argue that labs tune personality to manipulate users; the sourced position is that over-trust is possible, not intended.
Related guides
- Personal Model Benchmarks - score feel and capability on your own task set.
- When Labs Control the Model and Harness - what moves when the vendor owns the harness.
- Cost-Aware Model Routing for Agents - keep the persona-critical path pinned while commodity traffic moves.
- Agent Containment - bound what an assistant with tools can do, which matters more when persona is the surface under test.
Research basis: shared/abs-research-briefs/aidb/personality-is-ux/personality-is-ux-research-2026-09-23.md (verified 2026-09-25).
Sources
- The AI Daily Brief - Opus 5.5 vs GPT-6 Sol and Luna (2026-09-23, timecoded transcript)
- OpenAI Model Spec, version 2026-08-18
- OpenAI Model Spec, version 2025-02-12
- Anthropic - Claude's Character (2024-06-08)
- The Verge - OpenAI says GPT-5.1 is warmer with more personality options (2025-11-12)
- The Verge - ChatGPT is bringing back 4o as an option because people missed it (2025-08-08)
- Cheng et al. - ELEPHANT: Measuring and understanding social sycophancy in LLMs (arXiv:2505.13995)
- Cheng et al. - Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence (arXiv:2510.01395)
- LMSYS Org - Does style matter? Disentangling style and substance in Chatbot Arena (2024-08-28)
- Fang et al. - four-week randomized study, MIT Media Lab and OpenAI (arXiv:2503.17473)
- Personality Without Persons? Big Five inventories on LLMs (arXiv:2607.02325)
- OpenAI API docs - Deprecations and notice periods
- Claude platform docs - Model deprecations
- Microsoft Foundry Models lifecycle and support policy
- Amazon Bedrock - Model lifecycle
- Menlo Ventures - 2025 mid-year LLM market update



Submit a take
Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.