Ad
Skip to content

Grok 4.5 is so cheap compared to Fable 5 and GPT 5.5 that benchmark gaps may not matter much

Image description
Nano Banana Pro prompted by THE DECODER

Update –

  • Added Artificial Analysis Benchmark

Update, June 9, 2026:

Artificial Analysis ranks Grok 4.5 fourth on its Intelligence Index, behind Fable 5, GPT-5.5, and Opus 4.8. The model gained 16 points over Grok 4.3, putting SpaceXAI close to frontier performance. It trails OpenAI and Anthropic but beats all open-weights models and Google's Gemini.

Grok 4.5 is also very cost-efficient on the Intelligence Index, where a single task costs just $0.31. That's less than GLM-5.2 and Kimi K2.6, and five times cheaper than Claude Sonnet 5 (max), which scores lower on the same index. The independent AI benchmark service measures model performance across a range of tasks and aggregates the results into a single score.

Grok 4.5 is a big leap for xAI, even if it doesn't reach the top. And it's cheap. | Image: Artificial Analysis

According to Artificial Analysis, Grok 4.5 performs particularly well on agentic tasks. On the Coding Agent Index, Grok 4.5 running in Grok Build, xAI's equivalent to Claude Code, scores 76 points, matching GPT-5.5 in Codex and trailing Fable 5 in Claude Code by just one point, at a fraction of the cost. Per task, Grok 4.5 in Grok Build costs $2.49, compared to $5.07 for GPT-5.5 in Codex and $11.80 for Fable 5 in Claude Code. Grok 4.5 also averages just 1.9 million tokens per task, far less than GPT-5.5 (6.2M) and Fable 5 (7.2M).

Ad
DEC_D_Incontent-1

In its performance tier, Grok 4.5 is very cheap and efficient, assuming benchmark results hold up. | Image: Artificial Analysis

Artificial Analysis also flags a weakness. Accuracy on the AA-Omniscience Index rose from 35 to 52 percent, but the hallucination rate jumped from 25 to 54 percent too. The model knows more, but it's also more confident when it's wrong.

Original article, June 8, 2026:

xAI has released Grok 4.5. The model was trained on tens of thousands of Nvidia GB300 GPUs and targets coding, agentic tasks, and knowledge work.

Benchmark results paint a mixed picture. On Terminal Bench 2.1, which tests complex command-line tasks, Grok 4.5 scores 83.3%, nearly matching GPT 5.5 (83.4%) and trailing Anthropic's Fable 5 (84.3%) by just one point.

Ad
DEC_D_Incontent-2

But the gaps widen elsewhere. On DeepSWE 1.1, which measures the ability to resolve real GitHub issues, Grok 4.5 hits 53%, well behind OpenAI's GPT-5.5 at 67% and Fable 5 at 70%. On SWE Bench Pro, a curated set of harder software engineering problems, it scores 64.7%, beating Opus 4.8 (69.2% with max settings) in some configurations but falling short of Fable 5's 80.4%.

Model DeepSWE 1.1 Terminal Bench 2.1 SWE Bench Pro
Fable max 70% 84.3% 80.4%
GPT 5.5 xhigh 67% 83.4% 58.6%
Opus 4.8 max 59% 78.9% 69.2%
Grok 4.5 53% 83.3% 64.7%
GLM 5.2 44% 81.0% 62.1%

xAI says it relied on heavy data filtering, deduplication, and domain-specific selection during training to keep data quality high. The reinforcement learning stage covered hundreds of thousands of tasks, mostly from software engineering, with automated scoring. xAI built the training infrastructure for asynchronous learning, so agentic runs could stretch over many hours while training continued in parallel.

Grok 4.5 undercuts the competition on price

Grok 4.5 costs $2 per million input tokens and $6 per million output tokens. That's already far below the competition. Opus 4.8 runs $5 input and $25 output per million tokens. Fable 5 charges $10 input and $50 output per million tokens. GPT-5.5 and GPT-5.6 sit at $5 input and $30 output.

xAI also says Grok 4.5 uses 4.2 times fewer tokens than Opus 4.8 on SWE Bench Pro tasks and delivers results at 80 tokens per second. Lower per-token pricing and fewer tokens per task make Grok 4.5 by far the cheapest option in this performance tier, assuming the performance and efficiency gains hold up in practice.

Model Input (per 1M tokens) Output (per 1M tokens)
Grok 4.5 $2 $6
Opus 4.8 $5 $25
GPT-5.5 / GPT-5.6 $5 $30
Fable 5 $10 $50

The pricing strategy echoes what Chinese vendors like Zhipu and DeepSeek have been doing: get close enough on performance, then win on price.

Grok 4.5 is available now through Grok Build, Cursor, and the xAI console. Plugins are live for WordPowerPoint, and Excel. The model isn't available in the EU yet, with xAI targeting a mid-July launch. xAI trained Grok 4.5 alongside the code editor Cursor, which SpaceX acquired in mid-June for $60 billion in stock.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.