SpaceXAI

Grok 4.7

8.2/10
Maker
SpaceXAI
Origin
United States
Released
Sep 21, 2026

Strengths & weaknesses

  • SpaceXAI reports Terminal-Bench 4.0 nearly doubling to 38.0% from Grok 4.6's 20.3% on multi-hour terminal tasks
  • Holds Grok 4.6's $2/$6 per million tokens despite a larger base model and a longer reinforcement-learning run
  • 151st of 399 on LMArena's hard-prompts arena — the gains its maker reports on agentic benchmarks don't show up in blind human preference
  • SpaceXAI's own figures put it behind Claude Fable 5.1 on CursorBench 4.0, 46.3% against 51.8%

Evaluation

80% of weight measured

Scored on Grok 4.7 (xHigh) · sources as of Sep 26, 2026

Confidence band 8.2–8.5 · a model inside this range isn’t meaningfully apart from this one

Price & availability

Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.

w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works

More in Language Models

View all →
Anthropic
Under evaluation
  • Tops LMArena outright on hard prompts (1,545) and on the WebDev arena (1,827), the highest rating on either board
  • Anthropic reports roughly 40% lower cost per workload than Opus 5 and over 30% faster output, at $4/$20 per million tokens
  • Released weeks ago, so creative writing is still below this site's vote floor and no hallucination testing exists yet — its overall score is pending
  • Still the premium tier: many times the price of the fast models that most high-volume work runs on
OpenAI
8.8/10
  • 5th of 133 on LMArena's WebDev arena, OpenAI's strongest showing for building working web apps after Astra
  • 93.5% factual consistency on Vectara's hallucination leaderboard, ahead of GPT-6 Astra's 91.3%
  • 156th of 409 on LMArena's overall text arena — in blind comparisons people prefer a great many cheaper models
  • Sits below Astra in OpenAI's own line-up, so it doesn't carry the flagship's headline capabilities
OpenAI
  • Priced for volume at $0.10/$0.50 per million tokens, a twentieth of Sol, with the same 1.05M-token context
  • Built for high-volume clerical work — pulling fields from documents, summarising tickets — where throughput matters more than depth
  • 160th of 409 on LMArena's overall text arena and 155th on creative writing: not a model to reach for on hard or expressive work
  • OpenAI positions it two tiers below Astra, so complex reasoning and long agentic runs are outside its brief
Xiaomi
9.4/10
  • The strongest open-weight model measured here: 10th of 399 on LMArena's hard prompts, MIT-licensed at 1T parameters with 42B active
  • Native text, image, video and audio input with a 1M-token context, at $0.44/$0.87 per million tokens
  • Xiaomi is new to frontier releases, so its production track record is thin next to the established labs
  • 27th on creative writing against 10th on hard prompts — clearly stronger at reasoning than at expressive writing