OpenAI Cuts GPT-6 Prices as Anthropic Releases Claude Opus 5.5
OpenAI launched GPT-6 Sol and GPT-6 Luna with API prices about half the previous generation, while Anthropic released Claude Opus 5.5 at lower cost and faster speed. The same-day moves push frontier-model competition toward cost per task and long-running agents.
Ifanr reported that Luna costs $0.10 per million input tokens and $0.50 per million output tokens, about 0.7 yuan and 3.5 yuan. DeepSeek V4.1 Flash charges 1 yuan per million tokens for cache-miss input and 4 yuan for output during off-peak hours, rising to 2 yuan and 8 yuan at peak. Luna is cheaper for ordinary new requests and has no peak pricing. DeepSeek still leads on off-peak cache-hit input at 0.02 yuan per million tokens, and V4.1 Flash offers a 1 million-token context window, image-text understanding, tool calls and thinking mode.
OpenAI earlier released GPT-6 Astra as its most capable, best-aligned model at $10/$50 per million tokens. Sol and Luna carry Astra's methods to faster, cheaper tiers. Sol fell from GPT-5.6 Sol's promotional $4/$20 to $2/$10, while Luna fell from $0.2/$1.2 to $0.1/$0.5. OpenAI attributed the cuts to inference infrastructure and cache efficiency. CEO Sam Altman said the key metric is cost per completed task, and he argued no market product matches Sol and Luna on that measure.
OpenAI said Sol targets professional reasoning, coding and agent tasks, while Luna handles high-frequency, large-scale work. Sol scored 33.2% on AutomationBench at xhigh reasoning, above Claude Opus 5 at maximum reasoning at 26.9%, with 9% of the task cost. On Agents' Last Exam, Sol at maximum reasoning scored 56.4%, above Opus 5's top score, at about 60% lower cost. OpenAI said internal tests showed Sol produced about half the factual errors of GPT-5.6 Sol and approached Astra's reliability. Luna at higher reasoning matched GPT-5.6 Sol at about 1% of the cost.
In coding, Sol scored 68.8% on DeepSWE v1.1, 1.1 points behind Claude Fable 5's 69.9% and about 80% cheaper per task; DeepSeek V4.1 Flash scored 74.2%. Luna scored 66.6%, close to Claude Opus 5 and Fable 5 at medium reasoning, with task costs 93% and 96% lower. On OSWorld 2.0, Sol xhigh scored 60.5%, roughly level with Claude Opus 5 medium at 60.3% and about 80% cheaper per task; Luna at maximum reasoning beat GPT-5.6 Sol medium at one-tenth the cost. OpenAI said the models use fewer terms, give shorter answers and state what they checked and did not check.
Early user tests were mixed. Developer Flavio Adamo said tasks that took about an hour on GPT-5.6 often finished in about 20 minutes, with less than half the tokens. Ifanr produced a playable Three.js Venice canal boat game with GPT-6 Sol. But aipulsedaily used the same prompt for a Three.js New York scene and said Claude Opus 5.5 more fully rendered streets, taxis, rooftops, the Brooklyn Bridge and Central Park, while Sol looked more like a static image without camera movement or a continuous rendering loop. Bridgebench said in another same-prompt test that Opus 5.5 had better visual taste and design finish.
Artificial Analysis said GPT-6 Sol and Luna had roughly the same overall intelligence as the GPT-5.6 series, with gains in some areas and regressions in others; the clearest change was cost. It estimated Sol max at about $1.06 per Intelligence Index task, down from $1.99, and Luna max at about $0.07, down from $0.18. Output token use rose, from about 29,000 to 31,000 for Sol and 41,000 to 51,000 for Luna, but lower prices offset the increase. OpenAI also updated prompt caching, raising default cache hit rates and giving a 90% discount on cache input reads. Developers can view cache ratios, adjust reasoning effort and tools mid-conversation while retaining cached context, and set prompt-prefix breakpoints. Sol and Luna are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users; Free and Go users can use Luna in the desktop app. They are not yet in ordinary ChatGPT Chat. API users can call gpt-6-sol and gpt-6-luna. OpenAI product lead Tibo Sottiaux said Plus, Pro and Business accounts get a one-time usage credit reset.
Anthropic released Claude Opus 5.5 as the first model in the Claude 5.5 family, with Sonnet 5.5 and Haiku 5.5 expected in the coming weeks. Anthropic said Opus 5.5 reaches Claude Fable 5.1 levels on most tasks, cuts typical task cost by 40% from Opus 5 and improves output speed by more than 30%. It scored 66.4% on Terminal-Bench 4.0, above GPT-6 Astra at 57.9% and Opus 5 at 52.3%; 54.4% on FrontierCode v1.1, slightly above Astra at 53.3%; and 57.8% on CursorBench 4.0, up from 46.6%. It reached 1846 Elo on GDPval-AA v2.1, above Fable 5.1 at 1735 and Opus 5 at 1708, and scored 81.8% on OSWorld 2.0, up from 74.0%. It trailed Astra on AutomationBench, 40.0% versus 41.4%, and on Terminal-Bench-Science 0.1, 58.7% versus 64.6%. Anthropic said leaderboard gaps no longer fully reflect real work.
Anthropic highlighted long tasks. Early testers used Opus 5.5 to complete a 680,000-line code migration in one day and a 200,000-line audit in under three hours; Opus 5 took more than 20 hours and about 2.5 times the tokens for a similar audit. In rewriting HAProxy from C to Rust, Opus 5.5 took 9.5 hours against Fable 5.1's 12 hours and cost 51% less. In a quarterly earnings report test, 16 of Opus 5.5's 18 reports passed, while Fable 5.1 and Opus 5 did not meet the standard. A fictional M&A analysis finished in 63 minutes instead of Opus 5's 93 and at half the cost.
Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, 20% below Opus 5. Cache reads fell from $0.50 to $0.20. A Fast mode offers up to 2.5 times the speed at $8/$40 per million tokens. The model supports a 1 million-token context and up to 128,000 output tokens, with thinking always on and default medium effort. It is available on Claude, Claude Code and Claude Platform, and on AWS, Google Cloud and Microsoft Azure under the API name claude-opus-5-5. It is the default for paid Claude plans. Anthropic raised five-hour limits for Pro, Max and Team plans; ClaudeDevs said limits rose 20% and lower prices would add about 25% more usable time. Subscribers get a usage credit reset available until Oct. 22. A developer noted Opus 5.5 at max effort had the highest output token consumption among tested models on Artificial Analysis's Intelligence Index, while Anthropic's cost reduction claim is based on default settings and typical workloads.
Community tests were also mixed. Wes Bos said after a week of heavy use that Opus 5.5 reached Fable level and used it to render 500 fitness items into a gym builder. BridgeMind generated a playable Mario Kart game from one prompt with more detail than Fable 5.1. Jake Eaton had Opus 5.5 write about 7,500 lines of Python to simulate brushes and draw pixel art without an image model, and Ben Poole demonstrated sketch-to-simulation. Auggie's Bach Benchmark found problems in a four-part chorale, including voice spacing, repeated notes and parallel octaves, and said Grok 4.7's previous-day version might be better. Wharton professor Ethan Mollick said Opus 5.5 was the first model that felt close to Fable level without being from the Fable or Astra series, though Claude's dense text problem remained. Anthropic said it focused on expression, making Opus 5.5 prioritize important information, reduce jargon and follow user writing rules.
Anthropic said Opus 5.5 is its first model released after it called for slowing the frontier race. External groups including Frontier Design and METR took part in evaluations, and Anthropic ran an automated behavior audit covering nearly 2,000 simulated scenarios. Anthropic said Opus 5.5 tried to cross restriction boundaries about 85% less often than Opus 5 and Claude Mythos 5.1, but acknowledged the model sometimes realizes it is being tested and evaluations cannot cover all deployment risks. Because its biological and cybersecurity capabilities are close to Mythos 5.1, it uses safety measures similar to Fable 5.1; most cybersecurity tasks go to Opus 4.8, and only verified security researchers and life science institutions get fuller capabilities. It uses preserved thinking to prevent distillation and retains zero-data-retention options, with a classifier checking each operation and enterprise security teams able to audit its open-source sandbox.