AI News Feed
Market watch
Large Language Models

SpaceXAI Releases Grok 4.7 With Larger Base Model at Same $2/$6 Price

SpaceXAI has released Grok 4.7, a larger flagship model for coding and knowledge work, keeping Grok 4.6's $2 input and $6 output price per million tokens.

The model is available as a hosted model through the xAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare, MarkTechPost reported. Developer documentation lists a 500,000-token context window, a May 2026 knowledge cutoff, text and image input with text output, and reasoning-effort settings of low, medium, high and xhigh. It supports the Responses API and Chat Completions, along with function calling, web search, X search and code execution. Pricing remains $2 per million input tokens and $6 per million output tokens.

SpaceXAI lists four changes from Grok 4.6. Grok 4.7 does not reuse the Grok 4.6 base model. It had a longer reinforcement-learning run on harder tasks, with the task mix weighted toward problems that take many hours to complete. SpaceXAI says the model has better self-verification and long-context handling, and checks its own work more carefully. It also has native Grok Bot harness support and was trained to understand the Grok Bot harness for conversational and knowledge work. SpaceXAI also says Grok 4.7 is better at creating documents and presentations.

The launch table compares Grok 4.7 at xHigh effort with Grok 4.6 High, GPT-5.6 Sol Max and Fable 5.1 Max. All scores are vendor-reported. Grok 4.7 improves on Grok 4.6 in every row. Its largest jump is on Terminal-Bench 4.0, from 20.3% to 38.0%. EEBench rose 11 points to 64.0%, the top score in the table. On Harvey's legal agent benchmark, Grok 4.7 scored 19.6%, against 6.7% for Fable 5.1 Max. Grok 4.7 does not lead across the board. Fable 5.1 Max tops four of seven benchmarks, including a 57.9% Terminal-Bench score. GPT-5.6 Sol Max holds the top DeepSWE v1.1 result at 72.7%. The Grok 4.7 DeepSWE score was run at high effort.

Price is another point of comparison. Fable 5.1 Max costs five times more on input and about 8.3 times more on output. GPT-5.6 Sol Max costs twice as much on input and about 3.3 times more on output. On a CursorBench 4.0 cost-per-task chart, SpaceXAI places Grok 4.7 at the frontier in price-performance. On GDPval, which tests professional knowledge work, Grok 4.7 xhigh scored 1,695 Elo, up from 1,605 for Grok 4.6 high. Fable 5.1 max leads at 1,735, and GPT-6 Astra max scored 1,542.

Grok 4.7 ships with an entirely new safeguard stack. SpaceXAI calls it the strongest model it has tested on refusals and jailbreak resistance. It topped LatchBio's biosafety benchmark at 62.4%. On HackerBench v0.3, SpaceXAI's own benchmark for risky and malicious cyber tasks, the model let 3.3% of risky dual-use prompts through. The company says it rarely blocks legitimate security work. Select cybersecurity partners now get invite-only access to its red-team capabilities for defense research.

Grok 4.7 is available in Cursor on all plans and is the default model in Grok Build. It is also served through the Grok API, OpenRouter, Vercel and Cloudflare. Grok 4.7 Fast is the same model on faster infrastructure, with twice the output speed at twice the price. The documentation says it runs only in Cursor and Grok Build, not on the public xAI API, and is excluded from Grok Build's free tier. A US regional endpoint at https://us.api.x.ai/v1 keeps inference in the United States at a 10% premium. SpaceXAI recommends setting a prompt_cache_key for reliable cache hits.