A practical guide to what changed, the benchmarks that matter, the cost controls developers should not miss, and one hands-on demo worth building.
It is the middle child of the Claude family, and the one most people will actually use. It is quick, capable, cheap to run, and free to use for all users without any subscription.ย
In this article, we go over the latest iteration of the Claudeโs Sonnet family. We put it to test to see whether its agentic claims had any truth to them or not. And how a regular user of Claude app benefit with this free upgrade.

Sonnet 5.5 is now the default model for all users of Claude App. If you use Claude without a subscription, this is the model you are talking to. Opus 5.5 stays behind a paid plan, so for most people, Sonnet 5.5 is simply what Claude is. In short, the following improvements have been made:
Why this release matters Sonnet 5.5 is not a replacement for Opus 5.5 on the hardest open-ended work. It is the model to look at when the task is well-scoped, repeatable, tool-heavy, or latency-sensitive.

Anthropic positions Sonnet 5.5 as a fast low-cost complement to Opus 5.5. In the Claude apps, Medium effort is the default. On the Claude Platform, High is the default. That difference matters because effort changes latency, token use, and how much the model verifies its own work.ย
Anthropic does not publish Sonnet 5.5 parameter count, layer count, mixture-of-experts layout, or other internal model architecture details… which is expected for any proprietary model. Any article that gives those numbers is speculating. Its key features, are out in the open though:
low, medium, high, xhigh, and max. Adaptive thinking makes it so that you canโt disable effort/reasoning. ย For technical leaders, this is a useful architecture view: input context, reasoning budget, tool loop, verification behavior, and output. Those are the levers that determine reliability and cost in production.
One of the biggest advantages of Sonnet 5.5 is that you don’t need a paid Claude subscription to try it. You can access the model through Claude’s free tier, although free users have usage limits that reset every five hours.

claude-sonnet-5-5. API usage is billed separately based on token consumption.ย | Cost item | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input tokens | $2 / MTok | $4 / MTok |
| Output tokens | $10 / MTok | $20 / MTok |
| Cache write | $2.50 / MTok (5m); $4 / MTok (1h) | $5 / MTok (5m); $8 / MTok (1h) |
| Cache read | $0.20 / MTok | $0.20 / MTok |
The key change is cost per task, not cost per token. Sonnet 5.5 keeps Sonnet 5 pricing, but Anthropic says it often finishes with fewer tokens and fewer tool calls, which can lower the total bill by up to 30%.ย
Low or Medium for chat, fast iteration, and clearly scoped agent steps.ย Medium for well-specified coding and multi-step tool use.ย High for harder or longer coding work.ย Xhigh and Max for workloads where your own evaluations show a measurable gain.ย Here are some of the tasks/domains across which Sonnet 5.5 has delivered state-of-the-art performance:
A basic ‘Hello, Claude‘ example does not show why this model is interesting. A better demo is visual QA: give Sonnet 5.5 a screenshot of a broken web page and the page’s CSS, then ask it to diagnose the mismatch and produce the smallest safe fix.ย
For this test weโd be using a CSS file named styles.cssย containing the style code for this page:ย

Prompt:ย
“You are debugging this customer-success dashboard.ย
Compare the screenshot with the attached CSS and identify the visual issues.ย
For each issue:ย
Preserve the existing visual design and make the layout responsive.โย
Output:

Sonnet 5.5 completed the debugging task in around 10 seconds and did more than just rewrite CSS. It correctly mapped visual issues to specific rules, suggested minimal fixes, added verification steps, and clearly called out areas where it was uncertain. What stood out most was its ability to combine screenshot understanding with code-level reasoning. I would still verify the changes in a browser before production use, but for multimodal debugging and frontend QA, the response was fast, practical, and surprisingly precise.
Prompt: Make a modern slick and punchy video for a modern startup that works on Artificial Intelligence.ย
It succeeds because the prompt leaves room for creative interpretation while giving the model three strong anchors: modern, slick, and punchy, with AI/startup as the subject. If the result has strong pacing, clean motion graphics, confident typography, and avoids the usual generic โAI glowing brainโ bullshit, itโs a very strong output.

There is a significant jump over Sonnet 5 is large in agentic coding and computer use (10% -> 70%). Terminal-Bench 4.0 rises from 10.3% to 70.6%, CursorBench 4.0 moves from 34.1% to 55.5%, and OSWorld 2.1 moves from 57.0% to 80.1%. On GDPval-AA and AA-Briefcase, Sonnet 5.5 lands very close to Opus 5.5, which helps explain why Anthropic is positioning it for everyday knowledge work.
More effort is not always better. Anthropic reports that Sonnet 5.5 scored lower at Max than at Xhigh on FrontierCode because extra review sometimes caused timeouts or out-of-scope edits. In agentic systems, overthinking can be a real failure mode.ย

Benchmark score should not be your only selection criterion. Measure completion rate, tool-call count, latency, token usage, and how often a human has to repair the result.
Claude Sonnet 5.5 is compelling because the upgrade is practical. It is faster, uses fewer tokens on many tasks, is substantially stronger at agentic coding and visual work, and keeps the same per-token price as Sonnet 5. For teams building coding agents, visual QA systems, document workflows, or tool-using assistants, it is an obvious model to evaluate.
The important lesson is to evaluate the system, not just the model. Tune effort, preserve prompt-cache behavior, define verification, and control scope. Sonnet 5.5 can be very efficient when the task is clear. It can also spend extra time and tokens when you ask it to be maximally thorough. The best deployments will treat those controls as part of the application architecture.
Note: Some of the images used in this article have been sourced from the official Sonnet 5.5 release blog.
A. The per-token price is the same, but Anthropic says completed tasks can cost up to 30% less because Sonnet 5.5 often uses fewer tokens and tool calls.ย
A. Yes. Anthropic lists 1M tokens as the default context window and 128K tokens as the maximum output for a normal request.ย
A. For well-scoped agentic coding, Anthropic recommends starting at Medium and moving to High for harder or longer tasks. For general API usage, High is the platform default.ย
A. No. Anthropic reports at least one benchmark where Max scored lower than Xhigh because extra review caused timeouts or out-of-scope edits.ย
A. Not for every workload. Anthropic still positions Opus 5.5 as stronger for the hardest open-ended work that needs sustained judgment.ย