Grok 4.7's Benchmark Scores: Did It Top All Models?

Grok 4.7 scores 46 on AA benchmark, placing fourth, with gains in coding but token costs up.

22/09/2026 02:419 min read

SpaceXAI has launched Grok 4.7, and independent evaluations place it fourth among leading AI labs.

The benchmarking firm Artificial Analysis gave the model a score of 46 on its Intelligence Index, which is 2 points above Grok 4.6. Higher rankings belong to Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5.

Grok 4.7 Moves SpaceXAI Into the Top Four AI Labs

Before the launch, Elon Musk stated that the model would be based on a 2.1-trillion-parameter system and surpass all existing models. He also mentioned that the training process incorporated data from SpaceX.

Grok 4.7 is here.

It's a notable improvement over Grok 4.6 at the same price and speed. pic.twitter.com/H3OTBbXyvO

— SpaceXAI (@SpaceXAI) September 21, 2026

On AA-Briefcase, a benchmark for realistic professional tasks, Grok 4.7 achieved 1657 Elo. This represents a 111-point increase over Grok 4.6 and places the model near Claude Opus 5's performance.

In the GDPval-AA evaluation, the model scored 1695 Elo, 90 points higher than its predecessor. Coding test results also followed the same upward trend.

When used with Grok Build, its coding agent, Grok 4.7 scored 56 on the Coding Agent Index, 9 points above Grok 4.6.

The model has surpassed GPT-5.6 Sol and now trails only Anthropic's Claude Fable 5.1, GPT-6 Astra, and Opus 5.

"Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol," the post read.

Elsewhere, the model's progress was minimal. It gained 4.5 percentage points on Terminal-Bench 4.0 and 3 points on GDP.pdf. However, scores declined on the AA-LCR and AutomationBench-AA tests. Grok 4.5 had topped AutomationBench in July, supporting Musk's claim that it matched Claude Opus.

The Gains Come With a Token Bill

The rate card has not changed. Grok 4.7 remains priced at $2 per million input tokens and $6 per million output tokens. Cache hits stay discounted at $0.50, and the 500,000-token context window carries over from Grok 4.6.

However, the model does significantly more work per answer. Artificial Analysis measured roughly 81,000 output tokens per Index task, compared with 36,000 for Grok 4.6 and 27,000 for GPT-6 Astra. Thus, steady per-token pricing still results in a higher cost per task.

That token appetite falls on a division that has already reported a $1.26 billion loss. Meanwhile, Musk said on September 14 that Grok 4.8 would complete training within a week.

He also expects Grok 5 to be the model that achieves artificial general intelligence. The next release will test whether SpaceXAI can climb the rankings without additional compute costs.

Share to

Disclaimer: this article comes from third-party media and is provided for reference only. It does not constitute investment advice. Crypto and other financial products carry significant price volatility risk, so please make your own decisions carefully.

Related articles