DeepSeek

DeepSeek V4.1 Flash

9.2/10
Maker
DeepSeek
Origin
China
Released
Sep 10, 2026

Strengths & weaknesses

  • MIT-licensed and openly downloadable at 552B parameters with only 8B active, served by 27 independent providers
  • DeepSeek reports 74.2% on Deep-SWE v1.1, ahead of both Opus 5 (74.0%) and GPT-5.6 Sol (73.0%)
  • 36th of 409 on LMArena's overall text arena, well behind the frontier models it matches on published coding benchmarks
  • A Flash-tier model by design: DeepSeek's own Pro line is stronger on maths and long reasoning

Evaluation

80% of weight measured

Scored on DeepSeek V4.1 Flash (Max) · sources as of Sep 26, 2026

Confidence band 9.2–9.4 · a model inside this range isn’t meaningfully apart from this one

Price & availability

Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.

w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works

More in Language Models

View all →
Anthropic
Under evaluation
  • Tops LMArena outright on hard prompts (1,545) and on the WebDev arena (1,827), the highest rating on either board
  • Anthropic reports roughly 40% lower cost per workload than Opus 5 and over 30% faster output, at $4/$20 per million tokens
  • Released weeks ago, so creative writing is still below this site's vote floor and no hallucination testing exists yet — its overall score is pending
  • Still the premium tier: many times the price of the fast models that most high-volume work runs on
OpenAI
8.8/10
  • 5th of 133 on LMArena's WebDev arena, OpenAI's strongest showing for building working web apps after Astra
  • 93.5% factual consistency on Vectara's hallucination leaderboard, ahead of GPT-6 Astra's 91.3%
  • 156th of 409 on LMArena's overall text arena — in blind comparisons people prefer a great many cheaper models
  • Sits below Astra in OpenAI's own line-up, so it doesn't carry the flagship's headline capabilities
OpenAI
  • Priced for volume at $0.10/$0.50 per million tokens, a twentieth of Sol, with the same 1.05M-token context
  • Built for high-volume clerical work — pulling fields from documents, summarising tickets — where throughput matters more than depth
  • 160th of 409 on LMArena's overall text arena and 155th on creative writing: not a model to reach for on hard or expressive work
  • OpenAI positions it two tiers below Astra, so complex reasoning and long agentic runs are outside its brief
SpaceXAI
8.2/10
  • SpaceXAI reports Terminal-Bench 4.0 nearly doubling to 38.0% from Grok 4.6's 20.3% on multi-hour terminal tasks
  • Holds Grok 4.6's $2/$6 per million tokens despite a larger base model and a longer reinforcement-learning run
  • 151st of 399 on LMArena's hard-prompts arena — the gains its maker reports on agentic benchmarks don't show up in blind human preference
  • SpaceXAI's own figures put it behind Claude Fable 5.1 on CursorBench 4.0, 46.3% against 51.8%