Models / Coding & Agentic Gemini 3 Flash Gemini 3 Flash 7.0 /10
Same model as Gemini 3 Flash in Language Models, listed again here because it’s used as a coding model and scored on that rubric. It counts once in the directory’s totals.
Maker Google
Released Feb 5, 2026 Strengths & weaknesses + Resolves 75.8% of SWE-bench Verified issues on the standard open harness, the highest figure measured in this directory + Holds a million tokens of context at the cheaper Flash tier, so a large repository fits in one pass − Rated well below the Claude and GPT flagship tiers on the WebDev arena, where people judge working web apps − Its coding evidence is one harness run: nothing here measures how it behaves across a long agent session
Evaluation 100% of weight measured Scored on Gemini 3 Flash · sources as of Sep 26, 2026
Confidence band 7.0–7.0 · a model inside this range isn’t meaningfully apart from this one
Code qualityw45 7.6
Rule A 75.8% resolved · 75.8% ÷ 10 · SWE-bench Verified, % resolved (mini-SWE-agent harness) (Gemini 3 Flash) · Feb 17, 2026
Web developmentw40 5.1
Rule B 1,439 rating (95% 1,430–1,447) from 10,900 votes · Scale 1,050–1,800, 0 to 10 · LMArena WebDev arena (gemini-3-flash) · Sep 25, 2026
Codebase contextw15 10.0
Rule B 1,048,576 tokens · Scale 8,192–1,048,576 tokens, 0 to 10 · OpenRouter, context window · Sep 26, 2026
Price & availability Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.
w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works
+ 600B mixture-of-experts with 27B active and a 1M-token context, built for long-horizon agentic work + $1.00 input and $2.70 output per million tokens, with open weights promised for October 2026 − Generated roughly 1.7× the median output tokens on Artificial Analysis's index, so real cost runs well above the headline rate − A preview with no independent coding-benchmark result yet, so it carries no overall score here
+ Excellent terminal automation, git operations, and CI/CD debugging + Far more token-efficient than reasoning-heavy rivals on routine tasks − Needs detailed, unambiguous instructions; struggles with vague requests − Smaller context window than some rivals, a constraint on huge monorepos
+ Strong at inferring intent from vague prompts and architectural context + 1M-token context supports coherent multi-file, cross-repo refactors − Can use many more tokens than leaner coding models on routine work − Narrates its reasoning at length, which slows down quick tasks
+ Open-weight performance within striking distance of proprietary leaders + Free to self-host, appealing for cost-sensitive or air-gapped teams − Requires serious infrastructure to run at full size − Tooling and IDE integrations are less mature than Copilot, Cursor, or Codex