Methodology
How scoring works
Every kind of model is judged on what matters for its job. An image model and a transcription model shouldn’t share a scorecard, so each section has its own rubric.
A 0–10 scale
Each criterion is scored from 0 to 10. Scores reflect hands-on testing and public benchmarks at the time of assessment, and carry an assessment date.
Weighted by what matters
The overall score is the weighted mean of the criterion scores. Weights are relative within a rubric and shown as a share of the total.
No partial verdicts
An overall score is only published once at least 60% of a rubric’s weight has been scored. Until then the model is shown as in assessment.
0 of 100 models scored so far
Language Models
| Criterion | Weight |
|---|---|
Reasoning Handles hard, multi-step problems without losing the thread. | 25% |
Accuracy Factually reliable, with low hallucination on real tasks. | 20% |
Long context Keeps recall and coherence across very long inputs. | 15% |
Speed Fast time to first token and high throughput. | 10% |
Cost efficiency Quality delivered per dollar at typical usage. | 15% |
Deployability Licensing, self-hosting, data controls, and ecosystem support. | 15% |
Coding & Agentic
| Criterion | Weight |
|---|---|
Code quality Writes correct, idiomatic, maintainable code. | 30% |
Agentic autonomy Plans and completes multi-step tasks with little supervision. | 25% |
Workflow fit Fits IDEs, terminals, repos, and CI without friction. | 15% |
Speed Responsive enough for interactive, in-flow use. | 10% |
Cost efficiency Token and seat cost relative to work delivered. | 10% |
Privacy & control Data handling, self-hosting, and enterprise controls. | 10% |
Image Generation
| Criterion | Weight |
|---|---|
Image quality Aesthetic and photographic fidelity of outputs. | 30% |
Prompt adherence Follows complex prompts, counts, and composition. | 25% |
Text rendering Legible, correctly spelled text inside images. | 10% |
Editing & control Inpainting, references, style and character consistency. | 15% |
Commercial safety Clear licensing and indemnity for business use. | 10% |
Cost efficiency Price per usable image. | 10% |
Video Generation
| Criterion | Weight |
|---|---|
Visual fidelity Sharpness, lighting, and realism frame to frame. | 30% |
Motion coherence Plausible physics and temporal consistency. | 25% |
Native audio Synchronised sound, dialogue, and effects. | 10% |
Creative control Camera direction, references, and editing tools. | 15% |
Speed Generation time per clip. | 10% |
Cost efficiency Price per usable second of footage. | 10% |
Audio, Voice & Music
| Criterion | Weight |
|---|---|
Output quality Naturalness, fidelity, and musicality — or transcription accuracy. | 35% |
Control Steering of voice, emotion, structure, or style. | 20% |
Language coverage Breadth and quality across languages and accents. | 15% |
Latency Fast enough for real-time or high-volume use. | 10% |
Licensing & safety Commercial rights and consent safeguards. | 10% |
Cost efficiency Price per minute or per track. | 10% |
Search & Enterprise
| Criterion | Weight |
|---|---|
Answer grounding Accurate answers anchored in retrieved sources. | 30% |
Citations Clear, checkable references back to sources. | 20% |
Integrations Connectors to company data, clouds, and tools. | 20% |
Governance Permissions, audit, data residency, and admin controls. | 15% |
Cost efficiency Predictable pricing at organisational scale. | 15% |
Enterprise by Function
| Criterion | Weight |
|---|---|
Functional depth How well it covers the real workflows of its function. | 30% |
Integration Fits existing systems of record and data. | 20% |
Time to value Speed and effort of rollout and adoption. | 20% |
Scalability Holds up across headcount, regions, and volume. | 15% |
Total cost Licence plus implementation and upkeep. | 15% |