The same sentence costs more in Arabic
Every model is priced per token, and every price list is easy to compare. What no price list tells you is how many tokens your text actually becomes — and that depends on the script you write in. The same meaning can cost 429 tokens in Arabic on one model and 1483 on another, for identical work at an identical price per token. These are measured, not estimated.
Your bill is the price per token multiplied by the number of tokens. The first half is published everywhere; the second half is not published at all. On an Arabic workload the second half varies by 3.6× across the models measured here, which can outweigh any price difference between them.
One sentence, eighteen ways
The same words through every tokenizer, so the only thing changing between these numbers is which vocabulary is reading them. English lands on about twenty tokens almost everywhere. The Arabic runs from 13 tokens to 52 — 4.0× apart, for one sentence nobody edited in between. A model that learned the script keeps whole words; one that didn't spends several tokens per letter.
In the UAE, everyone is Emirati through their love for this land and their contributions to it.
في الإمارات الكل إماراتي، بحبه لهذه الأرض وعطائه لها.
ALLaM 7B
SDAIALLMVocabulary 64,000A box marked · is a token that isn't a character at all — a fragment of a UTF-8 sequence that means nothing on its own. They are shown rather than hidden, because a row of them is the finding.
Note that this Arabic is shorter than its English on the better tokenizers. Single sentences vary, and this one is unusually compact in Arabic — the figures measured across the whole corpus, where Arabic costs more on every tokenizer without exception, are in the table below.
Does a small model cost less?
A small model is the usual advice for cutting cost: it runs on your own hardware, it is cheaper per token, and for many tasks it is good enough. On an Arabic workload that advice is incomplete, because a small model is often given a small vocabulary — and a small vocabulary is paid for by whichever script it was not built around.
But small does not mean bad, and large does not mean safe. Across the models measured here the two groups cover almost exactly the same ground: large models run from 1.19× to 3.83×, small ones from 1.32× to 4.22×. The best small model beats all but one of the large ones. The worst large model is worse than most of the small ones.
The clearest proof is a pair from one family. These score identically, to the second decimal, because they share a tokenizer — so on Arabic the smaller one carries no penalty at all for being smaller:
Ask what vocabulary a model was given, not how many parameters it has. The first tells you what your Arabic will cost; the second tells you nothing about it.
Arabic and Chinese, against English
Tokens needed for the same 6 articles, relative to English. Below 1.00 means the language is cheaper than English on that model. Sorted by the Arabic figure, best first.
Measured corpus — 6 UN-hosted articles
| Model | Vocabulary | English | Arabic | Chinese |
|---|---|---|---|---|
| ALLaM 7BSDAIALLM | 64,000 | 362 | 1.19× | 3.87× |
| Falcon H1 34BTIILLM | 261,120 | 350 | 1.31× | 0.87× |
| BLOOMZ 560MBigScienceSLM | 250,680 | 356 | 1.32× | 0.83× |
| Aya 101CohereLLM | 250,100 | 463 | 1.35× | 0.80× |
| Jais family 13BInceptionLLM | 84,992 | 356 | 1.37× | 2.11× |
| GPT-4o / GPT-5 familyOpenAILLM | — | 352 | 1.62× | 1.16× |
| gpt-oss 20BOpenAILLM | 199,998 | 352 | 1.62× | 1.16× |
| Fanar 1 9BQCRILLM | 128,256 | 357 | 1.63× | 1.77× |
| Gemma 2 9BGoogleLLM | 256,000 | 350 | 1.67× | 1.01× |
| Gemma 2 2BGoogleSLM | 256,000 | 350 | 1.67× | 1.01× |
| DeepSeek V3DeepSeekLLM | 128,000 | 351 | 1.74× | 0.90× |
| Qwen3 8BAlibabaLLM | 151,643 | 354 | 1.81× | 0.92× |
| Qwen2.5 0.5BAlibabaSLM | 151,643 | 354 | 1.81× | 0.92× |
| Llama 3.xMetaLLM | 128,000 | 354 | 1.82× | 1.22× |
| AceGPT v2 8BFreedomIntelligenceLLM | 128,000 | 354 | 1.82× | 1.22× |
| GPT-4 / GPT-3.5OpenAILLM | — | 354 | 3.25× | 1.76× |
| Phi-3.5 miniMicrosoftSLM | 32,000 | 399 | 3.60× | 1.88× |
| Mistral 7BMistral AILLM | 32,768 | 369 | 3.83× | 1.54× |
| SmolLM2 1.7BHugging FaceSLM | 49,152 | 351 | 4.22× | 2.52× |
It's the vocabulary, not the size
A tokenizer has a fixed vocabulary, and everything outside it gets broken into pieces. Models with a large vocabulary have room for Arabic and Chinese words; models with a small one spend that room on English and pay for every other script by the letter. This tracks vocabulary size almost exactly — and not model size. BLOOMZ 560M has 250,680 tokens in its vocabulary and handles Arabic better than models many times its size, while Mistral 7B does not. Two models in the same family, one sixteen times the other, score identically because they share a tokenizer.
Monthly and yearly, on your own numbers
Set your workload and the price you actually pay. The token counts are measured; the price is yours, because the models with published prices and the models with published tokenizers are not the same list, and pairing them would be a guess.
Input tokens only, at 1.15 tokens per English word measured on this corpus. Output is priced separately and usually higher; the same multiplier applies to it.
Models that publish no tokenizer
A count can only be shown where the tokenizer is published. These are left out rather than estimated:
- ClaudeAnthropic publishes no tokenizer, in any form.
- GeminiGoogle publishes no Gemini tokenizer; counts come only from a metered API.
- ALLaMSaudi Arabia's national model publishes no tokenizer file at its address.
- Aya ExpanseGated. Aya 101, measured above, is a different and older model.
Where each vocabulary came from
Meta, Google and Inception gate their own repositories, so three of the rows above were read from republications rather than from the owner. A copy is only as good as its faithfulness, so each was checked against a second republisher with no connection to the first: same vocabulary size, and the same token identifiers for the same text, down to the number. Where two unrelated parties agree exactly, the copy is the original.
- ALLaM 7Bread from JasperV13/Yehia-7B-DPO-Reasoning-previewchecked against ALLaM-AI/ALLaM-7B-Instruct-preview tokenizer.model, byte-identical
- Gemma 2 9Bread from unsloth/gemma-2-9b-itchecked against rinna/gemma-2-baku-2b
- Gemma 2 2Bread from unsloth/gemma-2-2b-itchecked against rinna/gemma-2-baku-2b
- Llama 3.xread from NousResearch/Meta-Llama-3.1-8B-Instructchecked against unsloth/Llama-3.3-70B-Instruct
Method
Comparing tokenizers across languages requires parallel text carrying broadly the same meaning. This benchmark uses the same 6 articles of the Universal Declaration of Human Rights from the United Nations' English, Arabic and Chinese pages because they are stable, substantial and closely aligned. The subject is incidental: this is a tokenization benchmark, not an assessment of human-rights knowledge or policy.
The sentence shown split token by token is deliberately separate from the benchmark corpus. It is quoted from the President's site, which publishes it in Arabic and English, and is included only to make token splitting visible. It is unusually compact in Arabic, so using it for the ratios would bias the comparison. Every ratio and cost figure here comes from the declaration; the quote supplies only the split you can see.
Each available published tokenizer implementation is run over that text: OpenAI's rank files for the GPT tokenizers and tokenizer files from Hugging Face for the rest. Where an owner's repository is gated, the provenance section identifies the documented, independently checked republication used instead. The token counts are exact for this corpus; no rule of thumb is used.
Ratios are computed over the whole corpus rather than averaged across articles, so a long article counts for more than a short one — which is how a bill works.
The register is formal and legal. Ratios move somewhat with register, so treat these as the shape of the problem rather than a figure to quote to two decimals on your own text.
Corpus Universal Declaration of Human Rights · Measured Sep 28, 2026