Claude’s Intelligence Isn’t Worth $20 to Me Anymore
Claude can be brilliant at coding and writing. But in 2026, being brilliant is no longer enough to win my $20.

Search for a command to run...
Claude can be brilliant at coding and writing. But in 2026, being brilliant is no longer enough to win my $20.

Insightful post! Learned about Kimi and Qwen for the first time – really appreciate the detailed matrix. Thank you!
Product Observations is a series where I analyze real software products, not to criticize them, but to understand the engineering and UX decisions behind them. Every article starts with a real observation, explores the possible trade-offs, and proposes thoughtful improvements.
Nothing is broken. Nothing is slow. Yet one small interaction subtly interrupts the experience, and it highlights an important lesson about designing AI products.
Nothing is broken. Nothing is slow. Yet one small interaction subtly interrupts the experience, and it highlights an important lesson about designing AI products.

How I turned a visual showcase into a fast, CMS-powered professional platform for recruiters, search engines, and AI-assisted discovery

A long context window sounds like an obvious advantage for an AI agent. The agent can retain more search results, tool outputs, intermediate reasoning, and evidence. Give it enough context, and perhap

Building the frontend architecture, guided journey, and integrations for a new insurance purchasing feature

I like Claude. That is precisely why I have a problem with Claude Pro.
Claude can write exceptionally well. Its coding models are genuinely competitive. Claude Code is one of the strongest arguments Anthropic has made for putting an AI agent directly into a developer's workflow.
But every time I consider paying $20 per month for Claude Pro, I end up asking a slightly different question:
What exactly am I paying $20 for?
Not:
Is Claude intelligent?
It obviously is.
Not even:
Is Claude better than ChatGPT or Gemini at some tasks?
It can be.
My question is much more boring.
And much more important.
What useful work does my $20 actually buy?
That is where Claude starts becoming difficult for me to justify.
Before criticizing Claude, this needs to be clear.
The lazy version of this argument would be:
Claude charges $20 just for coding and writing.
That isn't true anymore.
Anthropic currently lists Claude Pro at $20 per month when billed monthly. Pro includes more usage as well as Claude Code, Claude Cowork, Claude Design, Claude Science, Research, unlimited projects and access to more Claude models.
So Anthropic has clearly been building an ecosystem around the model.
That deserves credit.
Claude Code extends Claude into software engineering workflows. Cowork brings similar agentic ideas into broader knowledge work. Research handles web investigation. Claude can create files, execute code, connect to external services and work with projects.
So my criticism is not that Anthropic gives you nothing for $20.
My criticism is that the economics of the $20 tier increasingly feel mismatched with how I want to use AI.
And the reason starts with Claude's greatest strength.
Its intelligence.
For the first phase of the Gen AI race, choosing an AI product largely meant choosing a model.
Which model writes better?
Which reasons better?
Which generates better code?
Which understands a complicated prompt?
Those differences still exist.
But something has changed.
The number of models capable of doing genuinely useful work has exploded.
GPT is good.
Gemini is good.
Claude is good.
Kimi is good.
Qwen is good.
And depending on the task, the ordering changes.
This matters because Claude doesn't need to become worse for its intelligence premium to become less valuable.
Its competitors only need to become good enough.
Consider Kimi K3.
Moonshot AI's own published evaluation does not show Kimi K3 universally defeating Claude. The results are mixed, which is actually more interesting.
In Moonshot's reported coding results, Kimi K3 scores 77.8 against Claude Fable 5's 76.8 on ProgramBench and 88.3 versus 88.0 on Terminal-Bench 2.1. On SWE-Marathon, Kimi reports 42.0 versus Claude Fable 5's 35.0. [Source]
Claude wins elsewhere: 70.0 versus Kimi's 67.5 on DeepSWE and 86.6 versus 81.2 on FrontierSWE. [Source]
These are vendor-published comparisons, different benchmarks use different harnesses, and Moonshot itself documents important evaluation conditions and fallbacks. They should not be treated as proof that Kimi is universally better than Claude.
But they demonstrate something more important for my argument:
Claude no longer gets to compete in a world where only Claude can produce frontier-quality coding work.
The gap is getting crowded.
And Kimi doesn't need to beat Claude everywhere.
It only needs to make Claude's remaining advantage small enough that I start asking what else my $20 buys.
Maybe.
Let's grant Claude the strongest version of the argument.
Suppose Claude is the best coding model for your particular workflow.
Not second best.
Not approximately equal.
The best.
If you spend six hours every day working inside Claude Code and it saves you even a small fraction of that time, $20 can be an absurdly good deal.
I wouldn't argue otherwise.
But that's a specialist value proposition.
My AI usage isn't only coding.
I research.
I analyze documents.
I write.
I debug.
I explore ideas.
I work with files.
I generate visuals.
I investigate technical questions.
I prototype.
Sometimes I need a powerful reasoning model.
Sometimes I need a fast, cheap model to perform a trivial transformation.
And sometimes I don't particularly care which model answers me because several of them can already do the job.
That's where my calculation changes.
I am not looking to rent the world's most impressive text box.
I am buying an AI workspace.
ChatGPT Plus is also $20 per month.
OpenAI currently includes capabilities such as advanced reasoning, file uploads and analysis, image generation, Deep Research, voice, custom GPTs and access to its coding agent, Codex, within its broader paid ecosystem. Limits still apply, and some agentic workloads use their own allowance or credits.
Google takes another route.
Gemini combines its models with Google's existing ecosystem and Deep Research. Its research system can work with Google Search and, when authorized, sources such as Gmail and Drive. Google AI plans also expand access to Gemini's models and features.
Claude has increasingly expanded in the same direction through Code, Cowork, Research, connectors and its other products.
Good.
That competition is exactly what I want.
But it also means I no longer think comparing subscriptions by asking which company has the smartest flagship model makes much sense.
My $20 isn't competing against another $20 model.
It's competing against another $20 system.
That distinction changes everything.
This is where my biggest problem with Claude Pro appears.
Anthropic says Claude's plans operate with rolling five-hour usage windows, with paid plans also having weekly limits.
More importantly, Claude chat, Claude Code and the rest of your activity draw from the same overall usage pool.
Usage isn't a simple fixed message counter. Long conversations, model choice, complexity and features affect how quickly capacity is consumed.
That's understandable.
Frontier inference is expensive.
Every provider needs limits somewhere.
My problem is what happens after I hit them.
Anthropic's current answer is essentially: wait for the limit to reset, upgrade, or on paid plans enable additional usage credits billed at standard API rates.
That makes perfect economic sense for Anthropic.
I'm less convinced it makes sense for my experience as a $20 subscriber.
Because the product has already demonstrated that it has models at different capability and cost levels.
Why should exhausting my premium allowance mean that my best continuation path is waiting, upgrading or starting to pay additional usage?
Why can't the system degrade gracefully?
Google documents an approach I prefer.
When a Google AI subscriber reaches the relevant Gemini usage limit, they can continue the conversation with Flash-Lite.
Read that again.
Continue the conversation.
The intelligence level drops.
The experience doesn't have to stop.
That is a product decision I appreciate far more than another benchmark victory.
A lot of my work doesn't need the strongest model available.
If I have exhausted my expensive reasoning quota and then ask:
Rewrite this sentence.
I don't need your flagship model.
If I ask:
Turn these notes into JSON.
I don't need your flagship model.
If I ask:
Explain this TypeScript error.
Maybe I don't need it there either.
Use the cheaper model.
Keep me working.
Save the expensive intelligence for the tasks that actually require it.
That is not merely a quota implementation.
It is graceful degradation.
Software engineers already design systems around this principle.
When one capability becomes unavailable, a resilient product should preserve as much useful functionality as possible instead of converting a partial resource constraint into a complete workflow interruption.
AI products should be judged the same way.
This is the question I keep returning to.
Imagine Claude is 10% better than another model at a task.
That advantage has value.
But now introduce availability.
If the slightly weaker model remains available for my work while the stronger one is temporarily unavailable within my subscription allowance, the comparison changes.
Theoretical intelligence isn't the same thing as practical utility.
A Ferrari is faster than a Toyota.
That becomes surprisingly irrelevant when the Ferrari is sitting in a locked garage.
Claude Pro does allow users to buy additional usage rather than wait, so this isn't an absolute inability to continue. But that creates another question:
What exactly was the $20 subscription buying me if my normal workflow regularly pushes me into additional metered usage?
For some professionals, the answer will still be obvious.
The productivity gain may dwarf the additional cost.
For me, it doesn't.
There is another uncomfortable part of this equation.
Claude isn't only competing with ChatGPT Plus and Google's paid AI plans.
It's competing with free Claude.
And free Kimi.
And free Qwen.
And free tiers from Gemini and ChatGPT.
The free market has become ridiculously capable.
That changes what a paid subscription needs to justify.
I'm willing to pay for AI.
I already do.
But paying merely to move from “very intelligent” to “slightly more intelligent” is becoming harder to justify when the free or cheaper alternative already crosses my quality threshold.
For a difficult coding problem, that difference may matter enormously.
For rewriting a paragraph?
Probably not.
For summarizing a document?
Probably not.
For brainstorming UI states?
Probably not.
For routine research assistance?
Maybe not.
For a complex repository-wide refactor?
Now we have a real comparison.
That's why “Which AI is smartest?” is becoming the wrong purchasing question.
The correct question is:
Where does additional intelligence materially change the outcome?
This is also why my own AI usage has become increasingly fragmented.
I use different systems for different jobs.
ChatGPT is where I currently pay for the broadest collection of tasks.
Gemini handles research and Google-centric workflows particularly well for me.
Claude remains useful when I want a second opinion on code, debugging or writing.
Other models and products enter when they provide something useful at a lower marginal cost.
That isn't brand loyalty.
It's routing.
And I increasingly think that is how technical users should think about AI.
Use expensive intelligence where expensive intelligence changes the result.
Use cheap intelligence where cheap intelligence is sufficient.
Use specialized tools where the surrounding workflow matters more than the underlying model.
Which leads to another uncomfortable question for Claude.
Suppose coding represents 90% of your AI usage.
In that case, I understand paying for Claude.
But I'd still compare Claude Pro against coding environments, not only ChatGPT and Gemini.
Tools such as Cursor, and Antigravity have made the model itself increasingly swappable. The important product becomes the environment around the models: repository context, editing, agentic execution, model choice and the development workflow.
Today Claude might be best for one task.
Tomorrow GPT might be.
Kimi might win another.
Gemini might be better somewhere else.
Why should my entire workflow depend on predicting which model family will remain ahead?
The more interchangeable frontier models become, the more attractive the layer above the model becomes.
No.
That would be too easy.
And wrong.
If Claude's particular strengths align with your work, $20 may be trivial compared with the value it creates.
If Claude Code saves a professional engineer one productive hour per month, arguing over twenty dollars starts looking silly.
If you prefer Claude's writing enough that alternatives genuinely reduce the quality of your work, the subscription may also make sense.
If Cowork becomes central to how you handle knowledge work, same conclusion.
Value is workload dependent.
But my workload isn't Claude shaped.
I need coding and research and files and visual generation and analysis and broad everyday assistance.
I also care about how a product behaves after I've consumed its expensive resources.
And increasingly, I care less about owning access to one company's smartest model than having access to a system that routes me toward the right capability for the job.
Under those conditions, Claude Pro becomes difficult for me to justify at $20.
Not because Claude isn't intelligent enough.
Almost the opposite.
Anthropic helped make model intelligence extraordinarily useful.
The rest of the industry responded.
Now the market contains proprietary and open models competing aggressively across coding, reasoning, multimodality and agents.
That means the distance between “best” and “good enough” can matter less to consumers than the distance between their surrounding products.
Claude can keep winning benchmarks.
It can keep becoming better at code.
It can keep becoming a better writer.
But if another system is already intelligent enough for my work and gives me more ways to use that intelligence, the benchmark lead stops deciding where my subscription goes.
Being smarter and being worth more are not the same thing.
And that's ultimately why Claude's intelligence isn't worth $20 to me anymore.
I don't need Claude to become worse before I cancel the subscription.
I just need everything else to become good enough.
And that is already happening.
This is probably the fairest way to explain why I struggle to justify Claude Pro.
I'm not against paying for AI.
I already pay for it.
My current AI stack looks roughly like this:
| Tool | What I pay | What I use it for |
|---|---|---|
| ChatGPT | $20/month | General chat, research, Deep Research, image generation, coding, and broader everyday AI work |
| Google One | ~$4/month | Research, custom research Gems, YouTube analysis, Google/Gmail-connected workflows, and general chat |
| Claude | Free | Smaller coding and debugging tasks |
| Kimi / Qwen | Free | General writing, presentations, and tasks where I don't need another paid model |
The exact prices aren't the important part here. Plans vary by region, promotions change, free tiers change, and AI companies seem determined to make pricing pages age faster than JavaScript frameworks.
What matters is how I allocate money.
I'm already willing to spend roughly $24 per month across AI products.
So Claude isn't competing against my unwillingness to pay $20.
Claude is competing for the next $20 in my AI budget.
And that's a much harder competition.
For $20, ChatGPT currently covers a broad portion of my workflow. I can move from an ordinary conversation to research, work with files, generate an image, investigate something more deeply, or move into coding without needing another subscription for each category.
My much cheaper Gemini subscription covers another useful part of my workflow, particularly research and the Google ecosystem.
Then there are tasks where I simply don't need premium intelligence.
A small debugging question? | Free Claude may be enough.
General writing? | Kimi or Qwen may already produce what I need.
A quick frontend prototype? | I can use Lovable.
That leaves me asking:
What additional $20 worth of problems would Claude Pro solve for me that this stack doesn't already solve?
And right now, I don't have a convincing answer.
This is an important distinction.
If I had no AI subscriptions at all, Claude Pro would be easier to evaluate.
Is Claude worth $20?
Maybe.
But that's not my actual purchasing decision.
My decision is:
Given the capabilities I already have, is adding Claude Pro worth another $20 every month?
That's a marginal value question.
And the answer changes once capabilities overlap.
Suppose Claude gives me excellent writing.
Useful, but I already have several models that write well enough for my needs.
Suppose Claude gives me excellent coding.
Much more interesting, but I already have coding capabilities through my existing paid tools, free Claude access, and increasingly capable alternatives such as Kimi.
Suppose Claude gives me Research, Cowork, artifacts, connectors and other agentic capabilities.
Again, useful.
But now I'm comparing those capabilities against research, agents, integrations, file workflows and other tools I'm already paying for elsewhere.
Every overlapping capability reduces the incremental value of adding another subscription.
Claude therefore doesn't need to prove that it is good.
It needs to prove something harder:
That the gap between what I already have and what Claude Pro adds is worth another $240 per year.
For my current workload, I don't think it does.
And this is where AI subscription comparisons often go wrong.
We compare:
Claude Pro $20
versus
ChatGPT Plus $20
versus
Google One $4
as though everyone begins with an empty toolbox.
Most power users don't.
We already have overlapping models, free tiers, developer tools, research products and specialized applications.
The economically relevant question isn't simply:
Which subscription gives me the most?
It's:
Which next subscription adds the most capability I don't already have?
For me, Claude currently loses that calculation.
Not because Claude is bad.
Because too much of what makes Claude good is already available elsewhere in my stack (i.eChatGPT Plus and Google One).