Methodology

How scoring works

Every kind of model is judged on what matters for its job. An image model and a transcription model shouldn’t share a scorecard, so each section has its own rubric.

A 0–10 scale

Each criterion is scored from 0 to 10. Scores reflect hands-on testing and public benchmarks at the time of assessment, and carry an assessment date.

Weighted by what matters

The overall score is the weighted mean of the criterion scores. Weights are relative within a rubric and shown as a share of the total.

No partial verdicts

An overall score is only published once at least 60% of a rubric’s weight has been scored. Until then the model is shown as in assessment.

0 of 100 models scored so far

Language Models

CriterionWeight
Reasoning
Handles hard, multi-step problems without losing the thread.
25%
Accuracy
Factually reliable, with low hallucination on real tasks.
20%
Long context
Keeps recall and coherence across very long inputs.
15%
Speed
Fast time to first token and high throughput.
10%
Cost efficiency
Quality delivered per dollar at typical usage.
15%
Deployability
Licensing, self-hosting, data controls, and ecosystem support.
15%

Coding & Agentic

CriterionWeight
Code quality
Writes correct, idiomatic, maintainable code.
30%
Agentic autonomy
Plans and completes multi-step tasks with little supervision.
25%
Workflow fit
Fits IDEs, terminals, repos, and CI without friction.
15%
Speed
Responsive enough for interactive, in-flow use.
10%
Cost efficiency
Token and seat cost relative to work delivered.
10%
Privacy & control
Data handling, self-hosting, and enterprise controls.
10%

Image Generation

CriterionWeight
Image quality
Aesthetic and photographic fidelity of outputs.
30%
Prompt adherence
Follows complex prompts, counts, and composition.
25%
Text rendering
Legible, correctly spelled text inside images.
10%
Editing & control
Inpainting, references, style and character consistency.
15%
Commercial safety
Clear licensing and indemnity for business use.
10%
Cost efficiency
Price per usable image.
10%

Video Generation

CriterionWeight
Visual fidelity
Sharpness, lighting, and realism frame to frame.
30%
Motion coherence
Plausible physics and temporal consistency.
25%
Native audio
Synchronised sound, dialogue, and effects.
10%
Creative control
Camera direction, references, and editing tools.
15%
Speed
Generation time per clip.
10%
Cost efficiency
Price per usable second of footage.
10%

Audio, Voice & Music

CriterionWeight
Output quality
Naturalness, fidelity, and musicality — or transcription accuracy.
35%
Control
Steering of voice, emotion, structure, or style.
20%
Language coverage
Breadth and quality across languages and accents.
15%
Latency
Fast enough for real-time or high-volume use.
10%
Licensing & safety
Commercial rights and consent safeguards.
10%
Cost efficiency
Price per minute or per track.
10%

Search & Enterprise

CriterionWeight
Answer grounding
Accurate answers anchored in retrieved sources.
30%
Citations
Clear, checkable references back to sources.
20%
Integrations
Connectors to company data, clouds, and tools.
20%
Governance
Permissions, audit, data residency, and admin controls.
15%
Cost efficiency
Predictable pricing at organisational scale.
15%

Enterprise by Function

CriterionWeight
Functional depth
How well it covers the real workflows of its function.
30%
Integration
Fits existing systems of record and data.
20%
Time to value
Speed and effort of rollout and adoption.
20%
Scalability
Holds up across headcount, regions, and volume.
15%
Total cost
Licence plus implementation and upkeep.
15%