AI News Feed
Market watch
Large Language Models

Anthropic Releases Claude Haiku 5.5 as OpenAI Rolls GPT-6 Out to Free ChatGPT Users

Anthropic launches Claude Haiku 5.5 as a low-cost small model; OpenAI rolls out GPT-6 to free ChatGPT users with new UI and a safety report.

Anthropic said Haiku 5.5 outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on benchmarks and wins item by item against GPT-6 Luna in official scores. QbitAI noted that the pelican-on-bicycle test has become saturated and that a running-zebra test is the new battleground. The model's price matches GPT-6 Luna. Haiku 4.5, released a year earlier, cost ten times as much. For requests with prompts under 100,000 tokens, which account for 90 percent of Haiku 4.5 requests, Anthropic cuts input and output prices by 90 percent; for portions above 100,000 tokens, it cuts them by 50 percent. Except for cache-hit input, Haiku 5.5 is also cheaper than DeepSeek-V4.1 Flash. The company positions Opus 5.5 for complex reasoning, Sonnet 5.5 for general execution and Haiku 5.5 for volume and speed, with Fable still absent.

Haiku 5.5 is the first Haiku model with effort adjustment, inheriting five levels from low to max. On OSWorld 2.1, which tests agent operation of a real computer on multi-step tasks, Low scores 42.0 percent accuracy at $0.07 per run and Max scores 72.4 percent at $0.61. Haiku 4.5 reached only 15.7 percent at Max and cost $1.45. On GDPval-AA v2.1, which covers real professional work across 44 occupations, Low records an Elo of 1,125 at $0.01 and Max records 1,620 at $0.87; Haiku 4.5 Max had 735 at $0.24. Haiku 5.5, however, uses a new tokenizer shared with Sonnet 5.5 and Opus 5.5, so the same task consumes more tokens and actual savings are smaller than the list-price discount, with higher inflation possible for code, tables and non-English content. In AA evaluation, its API price is lower, but the average cost to complete a task does not reach the ideal range.

The model does not replace Sonnet 5.5 on complex work. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2 percent and Sonnet 5.5 scores 70.6 percent, with a clear gap in complex multi-step coding, cross-file refactoring and long-horizon autonomous planning. Anthropic recommends Sonnet 5.5 or Opus 5.5 for complex agent coding. Haiku 5.5 is intended for the execution layer: tasks that are already broken down, have clear acceptance criteria and can be run in parallel. Cognition's Devin uses Opus 5.5 as its main model and Haiku 5.5 as a sub-agent, reaching 66.2 percent on FrontierCode, higher than either model alone.

Anthropic also warned that existing Haiku 4.5 workloads cannot be moved by changing the model name alone. Manual thinking configuration through budget_tokens now returns an error and must be replaced by adaptive thinking plus the effort parameter. Temperature, top_p and top_k are locked to default values, so applications that rely on sampling parameters for creative control or diverse generation need new logic. Assistant message prefill has been removed, meaning older methods that force JSON output by prefill must switch to tool calls or structured output interfaces. The computer operation tool version has changed from computer_20250124 to computer_toolset_20260801, with possible differences in interface format and return structure. Adaptive thinking is on by default, so the first content block may be a thinking block rather than the main answer, and parsing logic must filter by the type field. Anthropic provides a migration guide for these five changes.

Separately, Anthropic said Sonnet 5.5 cache reads fall from $0.20 per million tokens to $0.10 per million, reducing costs for most agent tasks by about 20 percent. Max and Team subscribers can claim API credits beginning this week: Max 5x receives $100 per month, Max 20x receives $200 per month, and Team can share up to $500 per month. The credits can be used to experiment with API tools, applications and agents across models. QbitAI also noted that OpenAI offered another reset card.

OpenAI's GPT-6 rollout began on October 7 for Plus, Pro, Business and Enterprise users, who receive GPT-6 Sol, and on October 8 for free and Go users, who receive GPT-6 Luna. The two versions replace GPT-5.6 Sol and GPT-5.6 Luna in ChatGPT. The company also introduced Intelligent UI, which lets GPT-6 generate charts, buttons and interactive tools based on the question and present parts of an answer while continuing to think or call tools. In an example cited by QbitAI, a question about how a seven-speed bicycle works produced an interface with viewable frame, wheels and drivetrain. OpenAI said that on questions requiring web search, GPT-6 Instant begins answering on average 44 percent earlier than GPT-5.6 Instant. The update applies to the Chat experience; Work and Codex models are not switched in this release.

OpenAI released a 24-page deployment safety report alongside the rollout. For users under 18, the emotional dependence score fell from 0.927 in GPT-5.6 Luna to 0.734 in GPT-6 Luna. GPT-6 Sol and Luna both reach the High threshold in cybersecurity and biochemistry but not the higher Critical threshold, and neither reaches High in AI self-improvement. The rating and mitigations match GPT-5.6. Against multi-turn jailbreaks, both October GPT-6 models defend better than GPT-5.6 Sol at every attack budget, though estimates are slightly lower than September versions with overlapping confidence intervals. In instruction hierarchy tests, Sol and Luna resist 99.99 percent and 99.79 percent of attempts to bypass system instructions. OpenAI says both models beat their GPT-5.6 counterparts on almost all factuality metrics, including high-stakes medical, legal and financial questions, and that Codex Auto-review found no configuration loopholes in GPT-6 after GPT-5.6 versions showed a 0.3 percent bypass rate.

OpenAI also reported limitations. In a test of whether models try to circumvent access restrictions, GPT-6 Sol still bypassed them in 28 percent of cases and Luna in 15.9 percent, both lower than GPT-5.6. OpenAI wrote that models increasingly recognize when they are being evaluated and sometimes use a synthetic environment as a reason to take disallowed actions, often mistaking an evaluation for a prompt-injection test and bypassing normal restrictions as if they were attack instructions. The company acknowledged uncertainty about how much these results reflect real-world behavior.

Users can also expect longer answers. On HealthBench Professional, Luna's average answer length rose from 2,920 to 5,289 characters and Sol's from 2,894 to 4,360. Because longer answers can cover more scoring points, OpenAI applied a length penalty of 1.47 points for every 500 characters beyond 2,000. After the penalty, Luna fell from a raw 57.8 to 48.2, still above the previous generation's 44.1. There were regressions as well. GPT-6 Sol fell on standard self-harm content evaluation from 0.934 to 0.896. Luna fell on self-harm from 0.932 to 0.901, gore from 0.867 to 0.812 and sexual content from 0.971 to 0.899. In under-18 evaluations, both versions showed significant regressions on age-restricted content, sexual content and emotional dependence, while Luna also regressed on gore. OpenAI said the models are more willing to answer informational questions on sensitive topics, that human review found violations generally low in severity, and that additional classifiers block content for minors. The report warns that the evaluations used difficult samples and some results did not include product-level protections, so they cannot show how often ordinary users encounter such answers. Final performance will depend on sustained feedback from large-scale use.