AI News Feed
Market watch
Large Language Models

Anthropic's Sonnet 5.5 Shows Agent Runtime Is the Core Change

Anthropic released Claude Sonnet 5.5 with unchanged API pricing, faster generation and higher Agent benchmark scores. Lei Feng News reports the bigger shift is in Agent runtime: fewer tool calls, more batching and lower iteration counts.

On the release page, Sonnet 5.5 still looks like a routine upgrade. API pricing is unchanged, generation is faster, and scores on Agent benchmarks such as Terminal-Bench and CursorBench are clearly higher. In Lovable's tests, Tool Call count fell by about one third and Shell Run nearly halved. In Base44's application construction tasks, average iterations dropped from 7.7 to 3.6. Anthropic also said the new model more frequently places multiple tool calls in the same batch.

Those numbers matter because Agent cost cannot be measured by tokens alone. A task also consumes resources through tool waiting, state rehydration, replanning, context growth and failure recovery. If the model-tool loop is long enough, those costs compound. Lei Feng News frames the engineering change as the model gaining static dependency analysis and execution graph compression, allowing a previously heavy Agent Runtime architecture to slim down.

An Agent must maintain a changing working state: user goals, verified assumptions, tool returns, code edits, environment state and unresolved issues. After each Tool Call, runtime merges new results into that state and asks the model to reassess the plan. The next reasoning round does not only read the latest test log. It needs to know why the test was run, which files changed, which assumptions were eliminated and how the workspace changed. Failure paths make this worse. If a coding Agent first blames authentication middleware, searches symbols, reads files, edits logic and runs tests, then discovers the issue is session serialization, the earlier code, edits, logs and intermediate judgments remain in context. A runtime that only appends history forces the correct path to work around residual state. A larger context window preserves history, but it does not tell the runtime which content still belongs in the current working set. Stable long-running Agents need to keep direct dependencies active, compress settled conclusions into structured state, and expire raw logs, repeated search results and observations no longer tied to dependencies.

Traditional ReAct execution is a strict serial chain: the model chooses one action, the tool executes, the result returns, the model judges again and triggers the next action. It is simple, but it assumes dependencies between every tool even when many operations could run together. In code troubleshooting, searching an exception string, reading package configuration, locating a test file and checking symbol references are often read-only views of the current code state and do not need to wait for one another. Batch Tool Call asks the model to make a local dependency judgment before entering the tool layer and group actions that can execute together. The real compression is the synchronization barrier. Instead of stopping after every result to restore state and replan, the model determines what information a stage needs, multiple barriers disappear, tools run in parallel, and results return in one batch. This requires partial-order planning: judging which actions have ordering, which only read state and which change state. Read-only operations are easier to parallelize. Write operations must consider read-after-write and write-after-write relationships. Tests must see modified code, caller changes may depend on an updated interface, and two subagents editing the same file can create write conflicts. Whether an execution graph can be compressed therefore depends less on how many Tool Calls are issued at once than on where synchronization points are placed.

Batch execution also creates heavy fan-in. A dozen concurrent tools may reduce waiting, but the next model round can receive large code, logs and search results. If these enter context raw, the saved synchronization cost returns as observation bloat. Runtime therefore needs result reduction after tools return: merge duplicates, shrink long logs into segments relevant to the current decision, filter low-value results and send compressed state back to the model. Sonnet 5.5 showing both more frequent batching and fewer total Tool Calls suggests the model is forming a relatively complete set of information needs earlier in some tasks, reducing later follow-up queries.

When the model helps decide how the execution graph unfolds, effort changes role. In single-turn tasks, higher effort mainly increases internal test-time compute. In Agent settings, internal reasoning directly changes external actions, so the same compute budget placed in different positions can produce different results. If the model performs an extra dependency check before editing code and finds an interface ordering relationship between two change-sets, it may avoid a conflict, a failed test and a rollback. That extra reasoning increases internal computation but reduces external execution. If evidence is already sufficient and the model keeps expanding analysis, the reverse occurs: it may start more code review, subagents and validation, turning extra compute into more execution nodes. On FrontierCode, Anthropic saw Max effort score lower than Xhigh. One explanation in the report is that high effort made the model call a code-review skill more frequently and split work across multiple subagents, with some branches timing out or producing edits beyond the task boundary. Effort changes the branch factor, and more internal budget can expand the execution graph.

Effort strategy should therefore be node-level rather than fixed for an entire task. Low-risk steps such as reading a repository or searching symbols can use lighter reasoning to generate a query set quickly. Before cross-file writes, schema changes or high-side-effect actions, reasoning can increase and compute can go into dependency checks and change-set planning. After code is written, entering tests quickly is more useful because real environment feedback is more effective than continued internal inference. If validation fails, reasoning can increase based on the new observation rather than rescanning confirmed parts. Checkpoints are also needed. If an Agent restarts from the beginning of a task after every failure, retry lengthens the execution chain. If runtime saves a stable state before key write operations, a failure can roll back only the latest change and continue from a trusted node. Coding Agents can use independent worktrees, temporary branches or sandboxes to isolate candidate changes, let subagents validate in separate environments, merge successful branches into the main workspace and discard failed branches. At this layer, Agent runtime is less a model-tool loop and more a dynamic controller: it maintains working state, decides synchronization points, controls fan-out, allocates more reasoning to high-risk nodes, saves checkpoints before writes, and decides whether to continue, roll back or commit based on validation.

Lei Feng News says the durable conclusion from Sonnet 5.5 may be that model capability is beginning to directly affect Agent runtime structure. If a model judges dependencies better before acting, runtime can reduce synchronization points. If observations are compressed in time, the working set does not keep growing with task length. If effort is placed at nodes with higher error costs, extra test-time compute can replace later failed execution instead of expanding the search space. Input and output tokens remain a base measure for Agents, but more useful indicators include how many model-tool barriers a task crosses, how many batch results are actually used by later decisions, how far a failure requires rollback, how large the active working set grows during execution, and how many subagent branches never commit. Sonnet 5.5 returns the discussion to an engineering question: how to place more computation where it eliminates subsequent execution cost, rather than letting an Agent use more reasoning to create a larger execution graph.