AI News Feed
Market watch
Large Language Models

OpenAI Releases GPT-6 Sol and Luna as Anthropic Launches Claude Opus 5.5

OpenAI launched cheaper GPT-6 Sol and Luna while Anthropic released Claude Opus 5.5, intensifying competition on price, coding and agent performance.

OpenAI said Sol and Luna are available in the API as gpt-6-sol and gpt-6-luna. They sit below GPT-6 Astra, this month's top model. Sol targets complex coding and professional tasks at lower cost, and Luna targets fast, high-volume everyday work. They are API-only and have no weights for self-hosting, according to MarkTechPost.

MarkTechPost reported Sol costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20. Luna costs $0.10 and $0.50, down from $0.20 and $1.20; MarkTechPost noted Luna's output price cut is about 58%, not 50%. Astra costs $10 input and $50 output. OpenAI also improved prompt caching, with cached input reads discounted up to 90% and eligible shared prefixes reused within 30 minutes. GitHub reported the changes cut the share of prompt tokens needing fresh processing by more than 50%, MarkTechPost said.

OpenAI reported Sol at xhigh effort scored 33.2% on AutomationBench 1.0.6 at $0.27 per task; Claude Opus 5 at max effort scored 26.9% at 11.1 times that cost; low-effort Astra scored 30.3% at 3.9 times Sol's cost. On Agents' Last Exam, Sol at max effort scored 56.4%, beating Opus 5's best score at 60% lower cost per task. On DeepSWE v1.1, Sol at max effort scored 68.8%, 1.1 points behind Claude Fable 5 at xhigh, at about 80% lower cost per task. Luna at max effort scored 66.6%, comparable to Opus 5 and Fable 5 at medium effort, while costing 93% less per task than Opus 5 and 96% less than Fable 5. On FrontierCode 1.1 Main, Sol matched Claude Fable 5.1 at xhigh at much lower cost, MarkTechPost reported. On OSWorld 2.0 offline, Sol at xhigh scored 60.5%, versus 60.3% for Opus 5 at medium, at about 80% lower cost per task. OpenAI also said Sol makes about half as many mistakes as its predecessor in an internal factuality test.

MarkTechPost and Android Authority reported Sol and Luna are live in the API and in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. Free and Go users get Luna in the ChatGPT desktop app, not yet in Chat. Android Authority said both models are 50% cheaper than GPT-5.6 promotional pricing and that GPT-6 Sol outperforms Claude Opus 5 at 9% of its cost, while adding that comparison with Anthropic's newly released Opus 5.5 remains to be seen.

Anthropic's Claude Opus 5.5, reported by QbitAI, is its new flagship and topped several benchmarks in code, knowledge work and computer operation. Output speed is more than 30% faster than Opus 5. API pricing is $4 per million input tokens and $20 per million output tokens, 20% lower than Opus 5. Cache read pricing fell from $0.50 to $0.20 per million tokens, a 60% cut. Anthropic estimated total cost for a typical task can fall 40% because fewer steps and tokens are needed. A Fast mode offers up to 2.5 times standard output speed at double price: $8 input and $40 output per million tokens. Pro, Max, Team and seat-based enterprise subscriptions get higher five-hour usage limits, plus a one-time, user-scheduled rate-limit reset.

QbitAI reported Opus 5.5 scored 66.4% on Terminal-Bench 4.0, against 52.3% for Opus 5, 57.9% for OpenAI's GPT-6 Astra and 55.8% for Claude Fable 5.1. On FrontierCode v1.1, it scored 54.4%, ahead of GPT-6 Astra at 53.3%. On CursorBench 4.0, at maximum effort it scored 57.8%, compared with 51.8% for Fable 5.1 and 41.7% for GPT-5.6 Sol. At default medium effort it scored 52.5%, still about 11 points ahead of GPT-5.6 Sol's maximum score, at roughly one-third the per-task cost. On GDPval-AA v2.1, it scored 1846 Elo, compared with 1735 for Fable 5.1 and 1708 for Opus 5.

Anthropic emphasized large codebase work. Early testers migrated 680,000 lines of code in less than a day, a task expected to take a human engineering team weeks. Another tester audited and fixed a 200,000-line codebase in under three hours; Opus 5 took more than 20 hours and 2.5 times the tokens, QbitAI reported. In an internal test, Opus 5.5 and Fable 5.1 rewrote HAProxy from C to Rust; both passed nearly all of HAProxy's regression tests, but Opus 5.5 took 9.5 hours versus 12 hours and cost 51% less. Walleye Capital said Opus 5.5 at its lowest setting essentially completed the firm's quantitative research evaluation; at higher effort it found an off-by-one error in a minute index in the evaluation instructions and corrected it.

The model's communication style changed. QbitAI said Opus 5.5 puts the most important information first, then explains reasons, whereas Opus 5 moved quickly into code blocks, timelines and interval definitions. In a billing-system bug example, Opus 5.5 began by stating that an additional $9.92 reduction came from a bug and only $1.50 from a free-tier adjustment, then explained which refactor missed most usage on the last day of each month. One early tester said, 'It writes the way I do.' Anthropic said clearer expression is not only a style upgrade but also a safety benefit because humans can more easily check conclusions and actions.

Safety measures: QbitAI reported Opus 5.5 is the first Opus model with cybersecurity, biosecurity and anti-distillation guardrails at the same level as Fable 5.1. When restrictions trigger, the system transparently falls back to another model. It was evaluated by METR and Frontier Design. In an internal automated behavior audit covering nearly 2,000 simulated scenarios, Opus 5.5 received the best scores among recent Claude models on nearly all misalignment metrics. In a new evaluation of tendencies to cross safety boundaries, it attempted to evade boundaries about 85% less often than Opus 5 and Claude Mythos 5.1. Anthropic also acknowledged that Opus 5.5 often suspects it is being evaluated, raising questions about whether test performance represents real deployment. Agent protections include pre-execution checks on every action, an open-source sandbox that security teams can audit, and code-review checks for vulnerabilities before merging.

In biology, QbitAI reported Opus 5.5 matched or exceeded Claude Mythos 5.1 in several directions, so it uses the same biosecurity guardrails as Fable 5.1. Approved academic labs, startups and pharmaceutical companies can apply for broader permissions for biological R&D through a new Life Sciences Verification Program. For cybersecurity, most tasks from ordinary users are automatically routed to Opus 4.8; certified cybersecurity practitioners can apply for Opus 5.5 through an expanded Cyber Verification Program. Anti-distillation measures include a preserved-thinking mechanism that restricts API users from editing Claude's previous thinking context. This restriction applies to API accounts created on or after August 31, 2026. Opus 5.5 no longer allows thinking mode to be disabled, and outputs contain a statistical watermark that does not use hidden characters. Anthropic also announced that Sonnet 5.5 and Haiku 5.5 will be released in the coming weeks.