Anthropic Releases Claude Sonnet 5.5 With 70.6% on Terminal-Bench 4.0 at Unchanged $2/$10 Pricing
Anthropic released Claude Sonnet 5.5, the second model in its Claude 5.5 family, claiming 70.6% on Terminal-Bench 4.0 and up to 30% lower cost per task while keeping $2/$10 list pricing.
Anthropic reports four main upgrades over Sonnet 5. Output generation is more than 30% faster, making it the fastest Sonnet to date. Cost per task is up to 30% lower because the model needs fewer tokens and tool calls. Writing is clearer, with early testers calling it a better collaboration partner. Vision and long-horizon work also improved, and Anthropic says it is the first Sonnet to beat Pokémon Red using only screenshots. The model has a 1 million-token context window, 128,000 maximum output tokens and a June 2026 reliable knowledge cutoff. Adaptive thinking is on by default, with effort running across five levels ranging from low through medium, high, xhigh and max.
Benchmark scores reported by Anthropic put Sonnet 5.5 at 70.6% on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5 and 66.4% for Opus 5.5 at Xhigh. On CursorBench 4.0 it scored 55.5%, about two points below Opus 5.5's 57.8%. On FrontierCode 1.1 it scored 52.1% at Xhigh and 46.2% at Max, while GPT-6 Sol scored 49.3%. On GDPval-AA v2.1 it scored 1844, compared with 1846 for Opus 5.5 and 1449 for Sonnet 5. On OSWorld 2.1 it scored 80.1% on computer use, close to Opus 5.5's 81.8%. On Humanity's Last Exam it scored 64.5% with tools, up from 54.9%. Anthropic said the Max result was lower than Xhigh because at Max the model more often ran multi-agent code review, sometimes causing timeouts or out-of-scope edits that FrontierCode penalizes. Anthropic also states that Opus 5.5 remains clearly stronger on complex, open-ended work.
List pricing is unchanged from Sonnet 5: $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 and cache writes $2.50 per million. That is half of Opus 5.5's $4/$20. According to the report, savings come from token efficiency rather than a price cut. Balyasny Asset Management measured about 121,000 tokens per answer versus 497,000 on Sonnet 5. Base44 reported 3.6 iterations per app build, where Opus 5 took 7.7. Zendesk processed tickets 20% faster. Effort defaults differ by surface: Claude Code and the Claude apps default to Medium, while the Claude Platform defaults to High.
In competitive terms, Sonnet 5.5 and GPT-6 Sol both list at $2 per million input tokens and $10 per million output tokens for prompts up to 272,000 tokens. Gemini 3.1 Pro Preview lists at $2/$12 for prompts up to 200,000 tokens, and Claude Opus 5.5 lists at $4/$20. Sonnet 5.5 offers a 1 million-token context window and 128,000 maximum output, matching Opus 5.5. GPT-6 Sol offers 1,050,000 context and 128,000 output, while Gemini 3.1 Pro Preview offers 1,048,576 context and 65,536 output. Sonnet 5.5's knowledge cutoff is June 2026; GPT-6 Sol's is April 20, 2026, and Gemini's was not listed. Sonnet 5.5 uses adaptive thinking with five effort levels, GPT-6 Sol offers six effort levels from none to max, Gemini supports thinking, and Opus 5.5 has adaptive thinking always on. Sonnet 5.5 and Opus 5.5 accept text and image inputs, GPT-6 Sol accepts text and image, and Gemini 3.1 Pro Preview accepts text, image, video, audio and PDF.
On FrontierCode 1.1, Sonnet 5.5 scored 52.1% at Xhigh, GPT-6 Sol 49.3%, Opus 5.5 54.4%, and Gemini was not reported. On GDPval-AA v2.1, Sonnet 5.5 scored 1844, GPT-6 Sol 1487, Opus 5.5 1846, and Gemini was not reported. On Chartography without tools, Sonnet 5.5 scored 61.6%, GPT-6 Sol 53.6%, Opus 5.5 64.4%, and Gemini was not reported. Sonnet 5.5, GPT-6 Sol and Opus 5.5 are generally available, while Gemini 3.1 Pro Preview is in preview. None are open weights. Benchmark scores were vendor-reported by Anthropic, while GDPval-AA runs were by Artificial Analysis, and prices were standard API list rates verified September 28, 2026. Anthropic pointed users to its official announcement, migration guide and system card for more detail.