AdvertisementAdvertisementAdvertisementAdvertisement
AI

SpaceXAI Unveils New Model Grok 4.7

9/22/2026, 04:21 PM • Evgenia Sliv

(edited: 09/22/2026)

SpaceXAI Unveils New Model Grok 4.7

SpaceXAI has unveiled Grok 4.7, a new version of the AI model for programming, knowledge work, and executing extended agent scenarios. Developers have increased the size of the base model and expanded reinforcement learning by adding tasks that can take several hours to solve. Grok 4.7 also received improved mechanisms for verifying its own results and support for the Grok Bot agent environment. In the CursorBench 4.0 test, the model scored 46.3% compared to 40.4% for Grok 4.6. On multi-hour AA Briefcase tasks, the result was 1657 Elo points, while the previous version scored 1546. In EEBench, the score increased from 53% to 64%, and in the Harvey Legal Agent Benchmark – from 15.8% to 19.6%. These results were provided by SpaceXAI itself, reflecting the developer company's testing.

Additional measurements were conducted by Artificial Analysis. In its Intelligence Index, Grok 4.7 scored 46 points, two points higher than Grok 4.6. When used together with Grok Build, the new model scored 56 points in the Coding Agent Index compared to 47 for its predecessor. In the ranking of models operating in their own agent environments, Grok 4.7 ranked fourth. However, the increase in results is accompanied by higher token consumption: in xhigh mode, the model generated about 81,000 output tokens on a complex task on average, compared to approximately 36,000 for Grok 4.6. The context window remained at 500,000 tokens. The model works with text and images, forming responses in text form. Four reasoning modes are available for managing computations: low, medium, high, and xhigh.

The cost of the standard version has not changed: for contexts up to 200,000 tokens, processing 1 million input tokens costs $2, cached – $0.50, and output – $6. For longer queries, rates double. Simultaneously, SpaceXAI launched Grok 4.7 Fast with doubled generation speed and twice the cost; it is currently only available in Cursor and Grok Build, without connection through the public API. Developers also reported a redesign of the model's security system. In its own HackerBench v0.3, Grok 4.7 missed 3.3% of dangerous requests, and in LatchBio biosecurity testing, it scored 62.4%. Thus, the new version combines higher scores in several tests with increased token consumption on complex scenarios, maintaining the previous basic tariff structure.

Popular news