AI News Feed
Market watch
Large Language Models

Anthropic Releases Claude Haiku 5.5 With Adaptive Reasoning and 1M-Token Context

Claude Haiku 5.5 lifts OSWorld 2.1 to 72.4%, cuts average cost about 75%, adds adaptive reasoning and a 1M-token context.

The largest gain appeared in computer operation. On the OSWorld 2.1 offline subset, Haiku 5.5 scored 72.4%, up from 15.7% for the previous generation; Sonnet 5.5 scored 83.9%. OSWorld tests a model's ability to complete long-chain tasks through computer interfaces, including understanding screens, executing actions and continuing after environment feedback. The report said these results reflect completion on a specific test set and do not directly represent success rates across all browser or desktop tasks, because dynamic pages, permission limits and abnormal states require separate testing in real environments.

Knowledge work and reasoning scores also rose. Haiku 5.5 scored 1,620 on GDPval-AA v2.1, compared with 735 for Haiku 4.5 and 1,840 for Sonnet 5.5; the evaluation covers actual tasks in 44 occupational categories, including material analysis, information integration and work-product generation. On AA-Briefcase v1.1, Haiku 5.5 rose from 614 to 1,578. On Humanity's Last Exam, it scored 45.9% without tools and 57.4% with tools, also above the previous generation.

Programming results were mixed. On the FrontierCode 1.1 main test, Haiku 5.5 scored 46.4%, close to Sonnet 5.5's 52.1%. On Terminal-Bench 4.0, however, Haiku 5.5 scored 39.2%, while Sonnet 5.5 reached 70.6%. Terminal-Bench involves complex multi-step tasks in command-line environments, where models often must read files, execute commands, handle errors and adjust based on test feedback, requiring sustained planning and execution-state maintenance. Anthropic said in its release materials that Sonnet 5.5 and Opus 5.5 remain better suited to complex agentic coding tasks, while Haiku 5.5 is mainly aimed at clearly scoped subtasks, summarization and context compaction.

Enterprise early tests offered additional data. Asana reported that in AI Teammates-related tests, Haiku 5.5 reduced task completion latency by more than 30% compared with its current model and improved single-turn inference speed by up to 2.5 times. HubSpot measured an average score of 92.8% over three runs in a simulated CRM environment. AlphaSense scored Haiku 5.5 at 0.84 in 400 document question-answering tests, up from 0.76 for the previous generation. The report said these results came from different companies' internal tests with different tasks and evaluation standards, so they are useful for specific business scenarios but cannot be used as a unified model ranking.

Haiku 5.5 is the first Haiku model to support adjustable reasoning intensity. Haiku 4.5 required developers to set a fixed internal reasoning budget in advance when using extended thinking, regardless of whether the task was simple information retrieval or complex analysis involving multiple conditions. Haiku 5.5 uses Adaptive Thinking, allowing the model to decide whether to use internal reasoning and how much computation to invest based on task content, while developers use Effort settings to control overall reasoning intensity. The report gave the example that extracting a date from a document usually does not need long analysis, while comparing clauses across multiple contract versions requires handling more conditions and interrelated information.

The API changes are practical. Haiku 5.5 no longer uses the previous method of manually specifying a fixed number of thinking tokens, so older reasoning configurations need adjustment. Internal reasoning also consumes the model's output-token budget, meaning programs that set output limits only by expected final-answer length may find that reasoning uses too much space and the main text cannot be generated completely. With adaptive reasoning enabled, responses may include separate thinking data blocks, so applications must distinguish internal reasoning from final output rather than assuming the returned content starts with the answer. For agents that continuously use tools, there are consistency requirements between reasoning data and conversation history; developers who modify existing history messages in later requests may affect the validity of the original reasoning state. Haiku 5.5 also adds restrictions on some sampling parameters, so applications that relied on those parameters to adjust output behavior need to check interface compatibility. These changes mainly affect projects already using Haiku 4.5 extended thinking, and migration requires rechecking reasoning budgets, response parsing and multi-turn message handling in addition to replacing the model version.

Haiku 5.5 supports a 1 million-token context window and 128,000-token output, allowing it to process longer documents, large code content and continuous tool records. Anthropic did not set a single low price for the entire context window. For prompts up to 100,000 tokens, the price is $0.10 per million input tokens and $0.50 per million output tokens. When a prompt exceeds 100,000 tokens, input rises to $0.50 and output to $2.50 per million tokens. Cache pricing is also tiered: $0.01 per million tokens for cache reads on shorter prompts and $0.05 for longer prompts. Anthropic said about 90% of Haiku 4.5 requests were within 100,000 tokens, so that share receives a larger price reduction under the new pricing, though actual cost also depends on changes in token usage. Haiku 5.5 updated its tokenizer, so the same content may correspond to a different number of tokens, requiring re-measurement even if input text is unchanged.

For long-running agents, context management affects cost further. A coding task may continuously produce file-search results, code-analysis records and test logs; if every call carries the full history, input length grows as the task proceeds. Prompt Caching can reduce the cost of reading repeated content, while Compaction can condense completed operations into a shorter task state and reduce the historical information that later calls must process. Anthropic listed context compaction as one of Haiku 5.5's main application scenarios. In deployment examples, Rogo reported using Haiku sub-agents to extract revenue data from corporate 10-K filings, with a larger model handling later presentation creation. Cognition used Haiku 5.5 as an auxiliary model in Devin Fusion; under a configuration led by Opus 5.5, FrontierCode reached 66.2%, with reported lower cost and latency. The release materials did not provide full call counts, token distributions and cost details, so the savings from adopting the same architecture in general business cannot be calculated from those examples.

Anthropic also cut the cache read price of Sonnet 5.5 by 50%, from $0.20 to $0.10 per million tokens. According to Anthropic's calculations, this reduces Sonnet 5.5's running cost in most agent tasks by about 20%. The adjustment only concerns cache read fees; standard input and output token prices did not fall by the same proportion. Claude Haiku 5.5 is available through the Claude API, AWS, Google Cloud and Microsoft Azure, and Anthropic has begun updating its Python and TypeScript SDKs to add beta support for computer operation and browser operation. From the published data, Haiku 5.5 is clearly above the previous generation in computer operation, knowledge work and multidisciplinary reasoning, but complex terminal programming remains far behind Sonnet 5.5. On price, shorter prompts gain a more obvious cost advantage, while prompts above 100,000 tokens fall under another tier. Existing Haiku 4.5 applications need migration work covering reasoning configuration, token calculation and response parsing. If they previously used Sonnet for some high-frequency tasks, they can compare Haiku 5.5's completion rate, response time and call cost on the same test sets. Anthropic has provided model capability, interface changes and pricing information, but actual business cost improvements still need validation on corresponding workloads, especially in multi-turn tool-calling scenarios where per-request price, retry counts after failures and context growth must be accounted for in a complete task.

Editor's Summary

Anthropic's Claude Haiku 5.5 substantially improves computer-use and knowledge-work benchmarks and introduces adaptive reasoning, a 1 million-token context window and tiered pricing, while complex terminal coding remains behind Sonnet 5.5. Its larger cost benefit applies mainly to prompts under 100,000 tokens, and existing Haiku 4.5 applications must revisit reasoning budgets, token accounting and response parsing. The model is available on major cloud platforms, and SDK updates add beta computer- and browser-operation support.