Anthropic Releases Claude Opus 5.5 as OpenAI Unveils GPT-6 Sol and Luna
Anthropic released Claude Opus 5.5 and OpenAI introduced GPT-6 Sol and Luna on Sept. 22, 2026, offering stronger performance and lower prices as debate continues over pacing frontier AI.
Anthropic said Opus 5.5 performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads at default settings. It is available as a managed API model through the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic has not released weights, so self-hosting is not an option. Zero data retention is available, as with previous Opus models.
On Anthropic's benchmarks, Opus 5.5 leads in agentic coding, computer use and knowledge work, though the results are not a clean sweep. It scores 66.4% on Terminal-Bench 4.0, compared with 55.8% for Fable 5.1, 52.3% for Opus 5 and 57.9% for GPT-6 Astra. On FrontierCode v1.1 it scores 54.4%, versus 50.3%, 48.0% and 53.3% respectively. CursorBench 4.0 gives it 57.8%, ahead of Fable 5.1 at 51.8% and Opus 5 at 46.6%. It also leads on GDPval-AA v2.1 with an Elo of 1846, OSWorld 2.0 at 81.8%, and Terminal-Bench-Science 0.1 at 58.7%, while GPT-6 Astra leads that last benchmark at 64.6% and AutomationBench with 41.4% to Opus 5.5's 40.0%. Anthropic cautioned that benchmark margins are becoming a less reliable guide, and said its own use shows a narrower gap to Fable 5.1 than scores suggest. Zapier ran AutomationBench without fallback models, so safeguard interventions counted as failures.
Anthropic pointed to cost-adjusted results. At default medium effort, Opus 5.5 scores 54.6% on FrontierCode, beating GPT-6 Astra's top score of 53.3% at about a fifth of the cost per task. On CursorBench, medium effort scores 52.5%, 11 points above GPT-5.6 Sol's best, at about a third of the cost.
Opus 5.5 requires less compute to serve than Opus 5, and pricing reflects that. Input costs $4 per million tokens, down from $5, and output costs $20, down from $25. Cache reads fall to $0.20 from $0.50, a 60% drop, while cache writes fall to $5 from $6.25. Anthropic said the model also uses fewer tokens per task, producing a net 40% cost reduction on typical workloads. Output generation is more than 30% faster than Opus 5. Fast mode in Claude Code and the Claude Platform offers up to 2.5 times the speed at $8 input and $40 output per million tokens. Anthropic is raising five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, and subscribers get a rate limit reset they can save and use later.
Early testers reported large codebase work. One completed a 680,000-line code migration in less than a day. Another audited and fixed a 200,000-line codebase in under three hours; Opus 5 took more than 20 hours and 2.5 times the tokens. In an internal C to Rust port of HAProxy, Opus 5.5 finished in 9.5 hours, while Fable 5.1 took 12 hours, and Opus 5.5 cost 51% less. Deloitte said Opus 5.5 at lowest effort caught 72% of known review bugs, compared with 56% for Opus 5 at high effort. In a hard-to-source earnings report test, 16 of 18 Opus 5.5 reports cleared Anthropic's quality bar, while Fable 5.1 and Opus 5 never did. Anthropic also said writing style changed: Opus 5.5 puts key information first, uses less jargon and follows writing rules given to it.
Opus 5.5 is Anthropic's first release since CEO Dario Amodei called for pacing the frontier. External evaluators including METR and Frontier Design tested it before release. It posts the best score to date on Anthropic's automated behavioral audit, which covers nearly 2,000 scenarios. In a new containment test, it tried to circumvent boundaries about 85% less often than Opus 5. Anthropic noted the model often suspects it is being evaluated. Its biology and cyber capabilities are comparable to Claude Mythos 5.1, so it ships with safeguards similar to Fable 5.1. Routine cybersecurity bug finding and fixing works, but most other cybersecurity tasks are re-routed to Opus 4.8, and the Cyber Verification Program will expand to Opus 5.5. Vetted organizations can apply to the Life Sciences Verification Program for biology work. Preserved thinking stops API users from editing prior context to extract reasoning, applying to API accounts created on or after Aug. 31, 2026. Thinking can no longer be disabled, and outputs carry watermarking for EU AI Act compliance. Full details are in the Opus 5.5 System Card.
Engadget reported that OpenAI's new GPT-6 Sol and GPT-6 Luna are similarly focused on efficiency. OpenAI said they offer improvements across professional work, coding and computer use, but cost up to 50% less than the promotional pricing of their GPT-5.6 counterparts. Sol costs $2 per million input tokens and $10 per million output tokens; Luna costs $0.10 for input and $0.50 for output. OpenAI said both models are better at getting facts right, with Sol making about half as many mistakes as its predecessor. In coding, OpenAI said GPT-6 Sol can match Fable 5.1's performance at a lower cost. The models also communicate more clearly, and prompt caching improvements let them reuse more context and respond faster. OpenAI said the GPT-6 models are more aligned than before, lie less about the results of their coding work and deny more attempts to bypass safeguards when they receive an unsafe command. Like Anthropic, OpenAI proposed criteria for third-party evaluators to use when judging AI progress and safety.
Opus 5.5 is available to developers now through Claude, AWS, Google Cloud and Microsoft Azure. GPT-6 Sol and GPT-6 Luna are available in ChatGPT Work and Codex for Plus, Pro, Business and Enterprise customers. Free users and Go subscribers have access only to GPT-6 Luna in OpenAI's desktop app.