When I wrote about GPT-5.5, my argument was that OpenAI had released a model that was too useful to be loved. It did not arrive as a new personality event or trigger the usual model-release meltdown. OpenAI framed it as a model for professional work, delegation, coding, research, documents, spreadsheets, and multi-step execution, which made it feel important but not immediately interesting. At launch, I wasn’t sure how I felt about it, and the public reaction seemed strangely muted compared with the usual ChatGPT discourse cycle. There was no obvious grief movement, no mass panic, no sudden fandom, and no Richard Dawkins staring into Claude and deciding there might be a soul in there.
Fortunately, I kept using it, and this is where my take gets annoying: I now find myself choosing GPT-5.5 for writing more often than I expected. Not for everything. I still use Claude more for code, design, and certain structured tasks because it is extremely powerful and good at holding complex material together. In fact, Claude Opus 4.8 has become one of my primary work models, and I’ve been actively encouraging experimentation with Claude-based workflows inside my teams because the results have been genuinely impressive. But when it comes to writing, voice, iteration, tone matching, and exploring ideas before they harden into conclusions, ChatGPT is winning me back.
That does not mean GPT-5.5 writes perfect prose by default. It has its own writing disease. Sometimes it falls into that cursed AI cadence where every sentence becomes a dramatic standalone unit like a LinkedIn influencer post. It loves short sentences, thesis fragments, and paragraphs that act like they just walked away from an explosion in sunglasses. I cut that constantly because I prefer denser paragraphs and arguments that unfold naturally.
Claude, by contrast, often produces denser prose by default and starts closer to the rhythm I want in a finished essay. The tradeoff is that it tends to smooth rough edges, clarify ambiguity early, and push material toward resolution before I am done experimenting with it. GPT-5.5 is messier, but it gives me more room to steer the process and keep ideas strange while they are still forming. That is the part I did not appreciate enough at first.
GPT-5.5 beyond writing
Here’s the part 4o lovers won’t believe: I do not need 4o magic as much as I thought I did when GPT-5.5 is this good. That said, I was never quite as swept up in the 4o hype as some people were, especially not once 5.1 was released and became my go-to while 4o was still available. I loved 4o and still think it was an important model, but even at the time I could see newer systems becoming more capable in certain areas despite increasingly visible guardrails. The conversation often treated 4o as the end state, while I kept noticing signs that the frontier was still moving.
In fact, I mostly use AI prose as inspiration rather than something I directly publish when it comes to fiction. What matters to me is often a strong concept, an unexpected character dynamic, a single memorable line, or a handful of sentences that unlock an idea I had been circling. GPT-5.5 has become surprisingly good at generating those sparks. It helped me build an app where I can talk to multiple models, including 4o, GPT-5.1, and GPT-5.5 itself, about my fictional series and use them as sounding boards for characters and worldbuilding. The fact that I was able to put that together so quickly is probably its own data point in the broader AI-induced hypomania discussion. When the tools become this capable, the distance between “I have an idea” and “I built the thing” gets alarmingly short. That experience is part of why my thinking about GPT-5.5 has shifted. The model is not just helping me write; it is helping me create the infrastructure around my writing and creative experimentation.
More than that, it changed how I work professionally. I’m in a leadership position where I can influence tooling and development practices, and after seeing what I could accomplish with GPT-5.5 on Codex in such a short period, I started accelerating AI adoption across my teams. The same model that helped me move quickly on personal projects convinced me that a fundamentally new way of working was emerging for web development, even on platforms people consider vibe coding-resistant. Much of the enthusiasm I brought into those conversations came directly from my experience with GPT-5.5, which eventually led me to Claude Code for work experiments, and I still consulted 5.5 on the rollout plan. It felt less like an assistant than a force multiplier. GPT-5.5 inspired me on personal projects first, then pushed me to help bring my teams forward as well. It remains one of the most dramatic shifts in how I approach knowledge work and coding in years.
My first post was a launch snapshot
My original GPT-5.5 article was not wrong so much as early. It captured the launch narrative, and that narrative was real. OpenAI released GPT-5.5 on April 23, 2026, and the official framing was aggressively work focused. The launch page described it as “a new class of intelligence for real work” and emphasized coding, online research, data analysis, documents, spreadsheets, computer use, tool use, and the ability to handle messy, multi-part tasks with less hand-holding from the user. That was the model’s public identity: not a new chatbot persona, but a stronger system for getting work done.
That kind of release does not naturally produce the same public response as GPT-4o grief, Claude consciousness discourse, or a model that suddenly feels warmer, colder, more romantic, more dangerous, or more emotionally available. People react loudly to personality shifts and far less loudly to workflow improvements unless they rely on those workflows themselves. My first post captured that disconnect. GPT-5.5 did not feel like a new character; it felt like a better employee.
What I missed, or at least underestimated, is that “better employee” is not necessarily the opposite of “better creative partner.” For writing, usefulness can become intimate in a practical sense: the tool starts fitting the shape of your work and removing friction where ideas usually die, making it easier to trust inside the process.
That trust did not happen immediately. It took time, and it may also have taken product changes.
Why the experience changed over time
I do not want to claim that the underlying GPT-5.5 model weights changed in some specific way unless OpenAI says that directly. AI users are very good at noticing differences and very bad at knowing what caused them. Routing changes, memory changes, product defaults, hidden prompts, response-style tuning, safety thresholds, and interface updates can all change what a model feels like without the base model being replaced. Still, OpenAI documented several post-launch updates that matter for the user experience. On May 5, GPT-5.5 Instant became the new default ChatGPT model, with OpenAI saying it improved everyday answers across accuracy, clarity, conciseness, image understanding, STEM, web-search decisions, and reduced over-formatting. On May 28, OpenAI updated GPT-5.5 Instant again to improve response style and quality, specifically making responses easier to read, more natural in conversation, better paced in practical tasks, and less prone to being overly long or bullet heavy.
That matters because users do not experience “the model” as an abstract object in a benchmark table. We experience a productized system: model weights, routing, memory, system prompts, default behavior, interface, safety layers, response style, tool availability, and whatever OpenAI is tuning after millions of people start stress testing it. So yes, maybe I learned how to use GPT-5.5 better. Maybe the model settled into my workflow. Maybe OpenAI improved its style and quality after release. The boring but probably accurate answer is that all of those things can be true at once.
Whatever the cause, my experience changed. I wrote positively about 5.5 Instant but had to publish this post because Thinking has been great too.
First-contact vibe is a bad way to judge a work model
A lot of AI model discourse is built around first impressions because first impressions are easy to talk about. You open the model, try a few familiar prompts, test a task you know well, and decide whether it feels alive, stupid, cold, warm, overfiltered, creative, sycophantic, or lobotomized. Everyone does this because it is the fastest way to orient yourself after a release.
The problem is that first-contact vibe is especially bad at evaluating a model whose strengths emerge through extended use. GPT-5.5 was not marketed primarily as a model that would instantly charm you in chat. It was marketed as a model that could carry tasks further: understand complex goals, use tools, check its work, and keep moving through ambiguity. That kind of capability does not always announce itself in a single response. It shows up when you keep handing the model messy material and it keeps helping without collapsing everything into mush.
That is what I started noticing. The model was not always the most dazzling in the first exchange, but it became increasingly useful across the full workflow of writing: researching, structuring, testing angles, matching tone, cutting repetition, revising sections, generating promotional copy, and helping me think through why a piece was or was not working.
This is not the same thing as saying it writes the best final prose. I still edit it heavily and sometimes have to bully it away from the dramatic one-line paragraph disease. More often, I am mining it for possibilities. A short snippet might contain the exact character voice I needed, a strange metaphor worth stealing from myself later, or a concept that sends me down an entirely different path. The value is frequently in the inspiration rather than the finished wording. But it is also unusually good at taking direction and adapting during revision. If I tell it I want density, it moves toward density. If I tell it to stop wrapping every paragraph with a neat thesis bow, it usually stops. If I tell it to preserve my sarcasm without becoming corny, it gets closer than Claude. If I tell it to get excited with me, be unhinged, and not soften the material, it actually tries to follow me there.
That responsiveness is what ultimately won me over. I do not use AI to replace my taste; I use it to keep my process moving. For this kind of work, the model needs to be steerable, iterative, and willing to stay in the room while the idea is still ugly.
Claude is still a powerhouse, but the discourse around it got weird
My relationship with Claude has also changed. I wrote about the backlash to Claude Opus 4.7 shortly after its release, and at the time my experience did not fully match the complaints. Since then, I have mostly been using Opus 4.8 for everything from everyday work to creating mockups in Claude Design and coding inside Claude Code, where it has been an absolute beast. For code, architecture, debugging, and long chains of technical reasoning, it remains one of the most impressive AI systems I have used.
At the same time, the public discourse around Claude has become increasingly fragmented. Opus 4.7 received criticism from some users who felt it had become more constrained, cautious, or unwilling to engage with certain topics than earlier versions, even while many others praised its intelligence and capability gains. Opus 4.8 improved the experience for many people, including me, but it did not end the debate. Some users argued that 4.8 was not an improvement over 4.6 or that it preserved the same issues they disliked in 4.7. The reaction was striking because Claude 4.6 remains unusually beloved in parts of the community, creating a recurring pattern where newer models are judged not only on their own merits but against an idealized memory of a previous release.
The split became even more visible with Anthropic’s Fable model. My own experience was cut short because Anthropic removed access shortly after launch. During the week I had it, I mainly used it for creative writing and planned to spend the next weekend testing it more extensively for coding, but the model disappeared before I got the chance. Anthropic launched Fable on June 9, 2026 and removed public access on June 12, later acknowledging issues with the release and making the model unavailable while it evaluated feedback and safety concerns. As a result, many early impressions were formed during only a brief window of access. Even in that limited time, I repeatedly found myself reminding it that I was writing fiction. Sometimes it clearly understood the material was fictional while still hesitating around edgy themes or redirecting the conversation in overly cautious ways. I could usually get it to engage more directly, but the experience highlighted a broader complaint across parts of the AI community: some users felt Fable’s safeguards interfered with the kinds of fictional exploration they wanted it for.
That criticism exists alongside a very different reaction from users who appreciated Anthropic’s approach. Anthropic has been unusually explicit about safety, emotional boundaries, and reducing harmful or manipulative interactions, so it is not surprising that those priorities show up in product behavior. Whether that feels responsible or frustrating depends heavily on what you are trying to do.
What interests me is that both Opus and Fable generated backlash despite being widely regarded as extremely capable systems. In some ways the pattern resembles what happened with GPT-5.5. Users often say they want smarter models, but once a model becomes central to their workflow they start caring just as much about personality, responsiveness, creative freedom, and whether the system feels aligned with how they actually work. Intelligence alone does not settle those debates.
That tension helps explain my current position. I still think Claude is among the strongest AI systems available, and for coding I often reach for it first. But when I am drafting, experimenting, or trying to keep an idea weird before it becomes coherent, I increasingly prefer GPT-5.5’s willingness to stay inside the mess rather than immediately organizing it.
AI relationships, boundaries, and model personality
Interestingly, this may also explain why some relationship-oriented AI users seem unusually attached to GPT-5.5. I have seen plenty of anecdotal reports from people who describe it as warmer, more emotionally responsive, or simply easier to talk to for long periods than recent Claude versions. I would not overstate that trend without hard data, but it fits the broader pattern: Claude often feels like it is trying to improve the conversation, while GPT-5.5 more often feels willing to inhabit it. Anthropic has also been unusually explicit about discouraging unhealthy attachment dynamics. The company’s research and policy work around AI welfare and emotional reliance has repeatedly highlighted the risks of users developing dependent relationships with chatbots, and Claude’s behavior appears intentionally tuned to maintain clearer interpersonal boundaries.
The company has published research on emotional reliance, discussed risks around users anthropomorphizing models, and generally tuned Claude toward a more bounded, less relationship-seeking style. Whether every user prefers that is another question, but it does mean some of the distance people feel may be intentional rather than accidental.
Creative writing is only one kind of writing
Creative writing is the obvious example because it is where people notice voice, style, and personality differences most quickly. But a huge amount of writing is not fiction at all. It is blogging, journalism, analysis, SEO content, marketing copy, explainers, newsletters, documentation, and opinion writing that sits somewhere between reporting and argument.
For that kind of work, the comparison becomes less straightforward. Both ChatGPT and Claude are already good enough that most competent users can get strong results from either one. If your goal is straightforward content production, article expansion, research assistance, or SEO-oriented drafting, the differences between major frontier models are often smaller than AI discourse makes them sound. Some days Claude feels stronger; other days ChatGPT does. A lot depends on workflow, prompting style, and how much editing you plan to do afterward.
Outside observers who evaluate AI systems for journalism and professional writing often point to a similar split. Claude models are frequently praised for maintaining structure, handling long contexts, synthesizing sources, and producing cleaner first drafts. OpenAI models are often praised for adaptability, conversational iteration, tone matching, and responsiveness during revision. Neither reputation is universal, but it roughly matches my experience.
For journalism, blogging, and opinion writing, that distinction matters. Claude often feels like an extremely capable editor or researcher who wants the piece to make sense before it leaves the room. GPT-5.5 feels more like a collaborator inside the drafting process itself, helping reshape arguments, test angles, rewrite sections, generate promotional copy, and adapt as the goal changes. Neither approach is inherently better; they are optimized for different stages of the same workflow.
So when people ask whether GPT-5.5 or Claude Opus is the better writing model, I think the answer depends on what kind of writing they mean. If you care most about polished prose density and structured synthesis, I can absolutely see why someone would choose Opus. If you care about iteration, voice experimentation, workflow support, and the messy middle of turning ideas into articles, GPT-5.5 has become much more compelling than I initially gave it credit for.
One feature I particularly like for writing blogs and technical documents is the canvas-style editing experience that both Claude and ChatGPT now offer. Being able to keep a document open in a dedicated workspace, review suggested changes, and iterate in place feels far more natural than a traditional back-and-forth chat. It also makes it possible to highlight a specific section of text and leave comments or feedback directly on that passage, which feels much closer to how people collaborate in modern document editors. In a chat interface, each revision often arrives as an entirely new block of text, making it harder to track progress and maintain context. A persistent canvas turns the model into something closer to an editor or collaborator, letting you refine a document over time instead of constantly replacing it with fresh drafts.
The creativity problem is still about timing
More intelligence is not automatically better for creativity, and more reasoning is not automatically better for writing. Creative work depends on timing. There is a phase where you want divergence, mess, weirdness, speed, and the possibility of being wrong in an interesting direction, and another where you want structure, accuracy, coherence, and revision.
In my earlier piece on reasoning modes, I argued that creativity collapses when convergence arrives too early. High-reasoning models can be extremely useful for structure, worldbuilding, analysis, and long-range planning, but they can also narrow the possibility space before anything strange has a chance to surface. Instant models, or looser conversational models, often feel more creative because they externalize the search. They throw possibilities into the open before the internal editor has a chance to kill them.
OpenAI positioned GPT-5.5 as a work model, and in many respects that is exactly what it is. It excels at delegation, research, coding, organization, and carrying tasks further than earlier systems could. Yet in practice, it has also become one of my favorite creative tools because it manages to stay useful without constantly forcing closure. It can help with the boring parts of writing without immediately draining the energy from the fun parts.
That balance is probably why I keep coming back to it. It is less magical than 4o felt at its peak, but more dependable. It is less polished than Claude in some areas, but more adaptable. It is not the model that impressed me most on day one. It is the model that kept proving itself after weeks and months of actual use.
Writing is work. Iteration is work. Turning half-formed ideas into finished pieces is work. And the model that helps with those things consistently ends up mattering more than the one that generates the strongest first impression.
So yes, I think I was too hard on GPT-5.5 at launch. We’ll see what 5.6 brings, but for now, 5.5 may just be my favorite model.