Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

Anthropic’s Claude Sonnet 5.5 Claims 30% Faster Output, Up to 30% Cheaper Costs

Anthropic's Claude Sonnet 5.5 beats Pokémon Red from screenshots, a cheaper, faster model with strong coding and design benchmarks.

By mitch·4 min read
A robotic arm holds a Game Boy cartridge glowing with a red screen in a modern laboratory.

Anthropic has released Claude Sonnet 5.5, a cheaper and quicker version of its large language model, and the company is making a bold assertion: Sonnet 5.5 is the first Sonnet model to defeat Pokémon Red using only screenshots. The announcement has drawn attention online, where people are weighing whether the model’s performance justifies the cost.

Sonnet 5.5 Compared to Sonnet 5

Opus 5.5 is positioned as the premium option, whereas Sonnet 5.5 is offered as a cheaper, quicker choice compared to it. Opus 5.5 is designed for handling demanding projects that call for careful deliberation, while Sonnet 5.5 excels at routine jobs, bug fixes, and producing polished documents, slides, and spreadsheets. The new model’s standout skill lies in its discerning eye for design.

Sonnet 5.5 carries over the same pricing from Sonnet 5, with $2 charged per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. The model usually requires far fewer tokens to deliver the same results, according to Anthropic, which says a task can cost up to 30% less than its predecessor.

Advertisement

Performance Benchmarks

On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scores 70.6%, compared to Sonnet 5’s 10.3%. On GDPval-AA, a test of real-world work across 44 occupations and nine major industries, Sonnet 5.5 scores nearly level with Opus 5.5 and about 400 points above Sonnet 5.

On FrontierCode, a benchmark for code generation, Sonnet 5.5 at High effort scores about one fifteenth of the cost per task while still earning 10 more points than Sonnet 5 at the same setting. On CursorBench, where models are tested on real coding tasks from actual Cursor sessions, Sonnet 5.5 comes close to Opus 5.5, falling within roughly two points of its best mark.

Safety and Alignment

Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5’s, so it launches with cyber safeguards and fallbacks like those Anthropic has developed for its most capable models. Its biology safeguards remain the same as Sonnet 5’s.

According to the company, Sonnet 5.5 marks the first Sonnet model to carry these protections. The two safeguards apply only to a limited set of high-risk requests; routine software development and most life sciences work remain untouched.

Sonnet 5.5 is the first Sonnet model to defeat Pokémon Red using only screenshots.

What Early Testers Saw

According to Anthropic’s own testing, Sonnet 5.5 runs at a lower cost than Sonnet 5. Early users observed that Sonnet 5.5 grasps a codebase rapidly and groups tool calls together more frequently than Sonnet 5, which reduces the number of steps and lowers expenses.

The team reported that it felt like a more natural conversational partner, and they noted its skill at design, saying it adds polish to user interfaces and follows slide templates to produce decks that need little editing afterward. During an internal test, the company provided it with a public company’s quarterly earnings materials and call transcripts alongside a slide template, then requested a 10-slide operating review. Two experts judged its initial draft to be ready to send as is.

The Pokémon Red Claim

What makes this announcement stand out is that it attaches a concrete figure to how Opus 5.5 stacks up against Sonnet 5.5. The assertion about Pokémon Red is presented as a bold statement, even if it has not been confirmed separately.

The company admits its own limitations. It says plainly that Opus 5.5 still does better on tasks demanding sustained judgment and open-ended work.

The Numbers

  • Terminal-Bench: Sonnet 5.5 scores 70.6%, vs. Sonnet 5’s 10.3%
  • GDPval-AA: Sonnet 5.5 nearly level with Opus 5.5, ~400 points above Sonnet 5
  • FrontierCode: +10 points vs. Sonnet 5 at High effort, at about one fifteenth the cost per task
  • Pricing: $2 per million input tokens, $10 per million output tokens, $0.20 per million tokens for cache reads

The Trade-Off

The compromise is simple: Sonnet 5.5 costs less per task because it requires fewer tokens to finish the same job. It also produces results more quickly than Sonnet 5, a claim made by Anthropic that positions it as the fastest Sonnet model yet.

At lower effort settings, Sonnet 5.5 costs less per task and performs comparably to Opus 5.5. At higher settings, it can perform comparably at a similar cost.

The Bottom Line

Anthropic’s own numbers show that Sonnet 5.5 improves on Sonnet 5. It runs faster, costs less, and handles design work better.

It remains to be seen whether the Pokémon Red screenshot win stands. The model does offer a cheaper route for ordinary tasks, though it does not entirely give up on quality either.

The notes themselves serve as a reminder of how far the field has advanced. The claim that a model can beat Pokémon Red from screenshots stands as a particular, bold assertion.

Source material: “Claude Sonnet 5.5,” Anthropic.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.