AI News Feed
Market watch
Large Language Models

SpaceX releases Grok 4.7, citing long-horizon processing gains and new safety safeguards

SpaceX introduced Grok 4.7, calling it its most capable large language model, with top results on coding, legal and safety benchmarks and pricing from $2 per million input tokens.

The Grok model series was originally created by xAI Corp., the startup founded by Elon Musk. Musk merged that company with xAI Inc. last year, and the combined organization was subsequently folded into SpaceX, which took on the Grok model line along with multiple artificial intelligence data centers, SiliconANGLE reported.

SpaceX evaluated Grok 4.7 on CursorBench 4.0, a benchmark built by Cursor, another company SpaceX recently acquired. Grok 4.7 completed the benchmark's challenges at an average cost of $4.69 per task, ahead of GPT-5.6 Sol and Fable 5.1. For both of those competing models, SpaceX tested hardware-intensive versions that prioritize output quality over cost-efficiency, meaning the comparison was not made at equal cost settings.

The company also ran Grok 4.7 against more widely used benchmarks. The model outperformed Fable 5.1 on the Harvey Legal Agent Benchmark and on EEBench, which contain legal and chip design tasks respectively. On EEBench, however, Grok 4.7 scored behind GPT-6 Astra, the latest large language model from OpenAI Group PBC.

SpaceX attributes the results to a new base model. Large language model training runs consist of multiple phases, each improving a different aspect of the algorithm under development; a base model is the initial version that emerges after the first phase, and AI providers frequently reuse base models across releases. SpaceX said its engineers also reworked Grok 4.7's reinforcement learning workflow, the training method used to strengthen a base model's reasoning. Compared with its predecessor, Grok 4.7 was given tougher training tasks and worked on them for longer stretches.

The model was built to work with SpaceX's Grok Bot harness, a collection of technical resources that lets Grok 4.7 divide complex work among multiple AI agents. Those agents can carry out different tasks in parallel, which speeds up processing, and check the accuracy of one another's output.

SpaceX also equipped Grok 4.7 with new safeguards. The company said the model set records on LatchBio and HackerBench, benchmarks that test a model's ability to block malicious biology research requests and malicious cybersecurity requests, respectively.

Grok 4.7 is priced from $2 per million input tokens and $6 per million output tokens. For latency-sensitive workloads, SpaceX offers a version that processes prompts twice as fast and costs twice as much.

The launch comes less than a week after SpaceX's previous model release. On Friday, the company introduced a text-to-speech model that it said offers double the accuracy of its predecessor at half the cost. SpaceX said Grok Voice Transcribe 2.0 also outperforms several competing entries in the text-to-speech category.

Editor's Summary

SpaceX released Grok 4.7, its most capable large language model to date, claiming a $4.69 average cost per task on CursorBench 4.0 ahead of GPT-5.6 Sol and Fable 5.1, and describing new multi-agent and safety features. The model trails OpenAI's GPT-6 Astra on EEBench, and is priced from $2 per million input tokens. It arrives days after the company's Grok Voice Transcribe 2.0 release, as SpaceX continues to consolidate artificial intelligence assets acquired along with the Grok series.