Anthropic has released Claude Sonnet 5.5, a cheaper and quicker version of its large language model, and the company is making a bold assertion: Sonnet 5.5 is the first Sonnet model to defeat Pokémon Red using only screenshots. The announcement has drawn attention online, where people are weighing whether the model’s performance justifies the cost.
Sonnet 5.5 Compared to Sonnet 5
Opus 5.5 is positioned as the premium option, whereas Sonnet 5.5 is offered as a cheaper, quicker choice compared to it. Opus 5.5 is designed for handling demanding projects that call for careful deliberation, while Sonnet 5.5 excels at routine jobs, bug fixes, and producing polished documents, slides, and spreadsheets. The new model’s standout skill lies in its discerning eye for design.
Sonnet 5.5 carries over the same pricing from Sonnet 5, with $2 charged per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. The model usually requires far fewer tokens to deliver the same results, according to Anthropic, which says a task can cost up to 30% less than its predecessor.
Performance Benchmarks
On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scores 70.6%, compared to Sonnet 5’s 10.3%. On GDPval-AA, a test of real-world work across 44 occupations and nine major industries, Sonnet 5.5 scores nearly level with Opus 5.5 and about 400 points above Sonnet 5.
On FrontierCode, a benchmark for code generation, Sonnet 5.5 at High effort scores about one fifteenth of the cost per task while still earning 10 more points than Sonnet 5 at the same setting. On CursorBench, where models are tested on real coding tasks from actual Cursor sessions, Sonnet 5.5 comes close to Opus 5.5, falling within roughly two points of its best mark.
Safety and Alignment
Sonnet 5.5’s cybersecurity capabilities are comparable to Opus 5’s, so it launches with cyber safeguards and fallbacks like those Anthropic has developed for its most capable models. Its biology safeguards remain the same as Sonnet 5’s.
According to the company, Sonnet 5.5 marks the first Sonnet model to carry these protections. The two safeguards apply only to a limited set of high-risk requests; routine software development and most life sciences work remain untouched.
Sonnet 5.5 is the first Sonnet model to defeat Pokémon Red using only screenshots.
What Early Testers Saw
According to Anthropic’s own testing, Sonnet 5.5 runs at a lower cost than Sonnet 5. Early users observed that Sonnet 5.5 grasps a codebase rapidly and groups tool calls together more frequently than Sonnet 5, which reduces the number of steps and lowers expenses.
The team reported that it felt like a more natural conversational partner, and they noted its skill at design, saying it adds polish to user interfaces and follows slide templates to produce decks that need little editing afterward. During an internal test, the company provided it with a public company’s quarterly earnings materials and call transcripts alongside a slide template, then requested a 10-slide operating review. Two experts judged its initial draft to be ready to send as is.
The Pokémon Red Claim
What makes this announcement stand out is that it attaches a concrete figure to how Opus 5.5 stacks up against Sonnet 5.5. The assertion about Pokémon Red is presented as a bold statement, even if it has not been confirmed separately.
The company admits its own limitations. It says plainly that Opus 5.5 still does better on tasks demanding sustained judgment and open-ended work.
The Numbers
- Terminal-Bench: Sonnet 5.5 scores 70.6%, vs. Sonnet 5’s 10.3%
- GDPval-AA: Sonnet 5.5 nearly level with Opus 5.5, ~400 points above Sonnet 5
- FrontierCode: +10 points vs. Sonnet 5 at High effort, at about one fifteenth the cost per task
- Pricing: $2 per million input tokens, $10 per million output tokens, $0.20 per million tokens for cache reads
The Trade-Off
The compromise is simple: Sonnet 5.5 costs less per task because it requires fewer tokens to finish the same job. It also produces results more quickly than Sonnet 5, a claim made by Anthropic that positions it as the fastest Sonnet model yet.
At lower effort settings, Sonnet 5.5 costs less per task and performs comparably to Opus 5.5. At higher settings, it can perform comparably at a similar cost.
The Bottom Line
Anthropic’s own numbers show that Sonnet 5.5 improves on Sonnet 5. It runs faster, costs less, and handles design work better.
It remains to be seen whether the Pokémon Red screenshot win stands. The model does offer a cheaper route for ordinary tasks, though it does not entirely give up on quality either.
The notes themselves serve as a reminder of how far the field has advanced. The claim that a model can beat Pokémon Red from screenshots stands as a particular, bold assertion.
Source material: “Claude Sonnet 5.5,” Anthropic.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

