Four Frontier Models in 30 Days Diverge on Price, Access and Cost per Task
Anthropic, OpenAI and Google DeepMind released four frontier models in 30 days. A MarkTechPost comparison finds similar benchmark ranges but wide differences in price, access and cost per task.
GPT-6 Astra and Claude Fable 5.1 share a list price of $10 per million input tokens and $50 per million output tokens. GPT-6.1 Sol and Gemini 4 Argon list at one-fifth of that, $2 and $10, though Argon's price is introductory and doubles later to $4 and $20. Access also differs. Astra is available through the OpenAI API, ChatGPT and Codex; Sol through the OpenAI API, ChatGPT Work and Codex; Argon only through the Fairwind Program; and Fable 5.1 through the Claude API, Bedrock, Google Cloud and Microsoft Foundry. Astra and Sol have 1.05 million-token context windows, while Fable 5.1 has a 1 million-token window and Argon's context window is not disclosed. Astra, Sol and Fable 5.1 cap output at 128,000 tokens per response. Argon's 1 million-token output cap is the only structural outlier.
Cached-input pricing matters most for agents that resend system prompts, tool schemas and history at every step. Astra's cached input costs $1.00 per million tokens, four times Fable 5.1's $0.25 and ten times Sol's $0.10. Argon's introductory cached-input rate is also $0.10. Astra charges $20 per million input and $75 per million output above 272,000 tokens, while Sol charges twice the input and 1.5 times the output above that threshold. Fable 5.1's long-prompt pricing is flat to 1 million tokens, and Argon's is not disclosed. None of the four models has open weights.
Benchmarks show no model sweeping the board. Google DeepMind's published comparison, which MarkTechPost described as vendor-reported, places Gemini 4 Argon ahead on long-horizon software engineering and knowledge work. Argon scored 77.9% on DeepSWE v1.1, against 74.1% for GPT-6 Astra and 67.4% for Claude Fable 5.1; 68.9% on the Vals Index, against 63.1% for Astra and 65.8% for Fable 5.1. Astra led FrontierSWE v2 at 65.5%, against Argon's 55.0% and Fable 5.1's 56.3%, and OSWorld-2.0 at 72.6% to Argon's 69.2%. On Terminal-Bench 4.0, Astra scored 58.2%, Argon 57.4% and Fable 5.1 57.9%; Claude Opus 5.5, which is not in this lineup, led that test. Argon and Astra tied at 68% on CWE-bench v1, with Fable 5.1 at 58%.
GPT-6.1 Sol does not appear in Google's table. OpenAI's own numbers place it close to Astra on several measures. Sol matches Astra on DeepSWE v1.1 at roughly one-fifth the cost, lands within 2.1 points of Astra on the OSWorld 2.0 offline set at about one-seventh the cost per task, and scores 2.2 points above Claude Opus 5.5 on AutomationBench 1.0.6 at medium effort, according to OpenAI. Astra still leads Terminal-Bench Science 0.1 at 68.1%, and OpenAI recommends Astra for the hardest research. Independent signals differ on raw intelligence. On the Artificial Analysis Intelligence Index, Astra scores 61 and Fable 5.1 scores five points higher; on the coding-agent index, Fable 5.1 in Claude Code scores 70 against Astra's 67. Artificial Analysis also reports that Argon equals Astra on the Intelligence Index. On ARC-AGI-2, Astra scores 95% and Fable 5.1 scores 90%.
Cost per task can separate models even when list prices are identical. Artificial Analysis puts Claude Fable 5.1 at $9.18 per task against $4.72 for GPT-6 Astra, about 1.9 times as much at the same list price. The gap comes from token volume rather than rates; Anthropic also says its newer tokenizer produces roughly 30% more tokens for the same text. For the cheaper models, Artificial Analysis reports that Argon equals Astra's Intelligence Index at 60% of Astra's cost per task using introductory prices. On Terminal-Bench Science, OpenAI reports $5.47 per task for Sol against $23.80 for Astra. Caching can reverse the ranking for agents. For a 200,000-token cached context before output tokens, Astra costs $0.20 per step and $20.00 per 100 steps, Fable 5.1 costs $0.05 and $5.00, and Sol and introductory Argon each cost $0.02 and $2.00. Astra's 200,000-token context remains under its 272,000-token long-prompt threshold. In cache-heavy loops, Fable 5.1 reads context at a quarter of Astra's rate. MarkTechPost's comparison concludes that buyers should measure both on their own traces before choosing on per-task headlines.