Skip to content
KoishiAI
ไทย

Thai Token Counter

Paste text to see how many tokens it takes in each model, and what it costs at today’s prices.

Your text never leaves your machine — everything is counted in the browser Prices from OpenRouter as of 2026-09-23
Try:
characters
Thai letters

Token count

  • GPT-4o · 4.1 · 5.x · 6
    o200k_base tokenizer · exact
  • Typhoon 2.5 · Qwen3
    ~2 MB, loaded on demand · exact
  • Qwen3.5 – 3.8
    ~3.5 MB, loaded on demand · exact
  • Claude Opus · Sonnet · Fable 5.x
    tokenizer from 4.7 on · estimate
  • Claude Haiku 4.5
    pre-4.7 tokenizer · estimate
  • Gemini 2.5 · 3.x
    2.5, 3.1 and 3.8 measured identical · estimate
  • Grok 4.6 · 4.7
    4.6 and 4.7 count the same · estimate
  • DeepSeek V4 · V4.1
    Flash and Pro count the same · estimate
  • Llama 4
    Maverick and Scout identical · estimate
  • Llama 3.3
    measured via OpenRouter · estimate
  • Mistral Medium 3.5
    measured via OpenRouter · estimate
  • Kimi K3
    measured via OpenRouter · estimate
  • Kimi K2.5 · K2.6
    measured via OpenRouter · estimate
  • GLM-5.3
    measured via OpenRouter · estimate
  • MiniMax M2.7 · M3
    M2.7 and M3 count the same · estimate
  • Amazon Nova 2
    measured via OpenRouter · estimate

Cost

Model Price in / out per 1M tokens Total
GPT-6 Astra $10.00 / $50.00
GPT-6 Sol $2.00 / $10.00
GPT-6 Luna $0.100 / $0.500
GPT-5.6 Sol $2.00 / $10.00
GPT-5.6 Luna $0.200 / $1.20
GPT-5.5 $5.00 / $30.00
GPT-5.4 mini $0.750 / $4.50
GPT-5 nano $0.050 / $0.400
GPT-4.1 mini $0.400 / $1.60
GPT-4o mini $0.150 / $0.600
Claude Opus 5.5 $4.00 / $20.00
Claude Fable 5.1 $10.00 / $50.00
Claude Sonnet 5 $2.00 / $10.00
Claude Haiku 4.5 $1.00 / $5.00
Gemini 3.1 Pro $2.00 / $12.00
Gemini 3.8 Flash $0.750 / $3.75
Gemini 2.5 Flash-Lite $0.100 / $0.400
Qwen3.6 35B-A3B $0.150 / $1.00
Qwen3 235B-A22B $0.087 / $0.350
Qwen3 30B-A3B $0.048 / $0.193
Qwen3.8 Max $2.00 / $6.00
Qwen3.8 Flash $0.150 / $0.470
Qwen3.7 Flash $0.030 / $0.130
Grok 4.7 $1.60 / $4.80
DeepSeek V4.1 Flash $0.100 / $0.500
Llama 4 Maverick $0.188 / $0.652
Llama 4 Scout $0.100 / $0.300
Mistral Medium 3.5 $1.50 / $7.50
Kimi K3 $3.00 / $15.00
GLM-5.3 $0.840 / $2.64
GLM-5.3 Flash $0.150 / $0.500
MiniMax M3 $0.300 / $1.20
Amazon Nova 2 Lite $0.300 / $2.50
GPT-5.6 Terra $2.00 / $12.00
Qwen3.7 Max $1.48 / $4.42
Qwen3.7 Plus $0.320 / $1.28
Qwen3.8 27B $0.420 / $3.00
Gemini 3.7 Flash $0.750 / $3.75
Gemini 3.5 Flash-Lite $0.300 / $2.50
Claude Opus 5 $5.00 / $25.00
Claude Fable 5 $10.00 / $50.00
Grok 4.6 $2.00 / $6.00
DeepSeek V4 Pro $0.953 / $1.91
DeepSeek V4 Flash $0.082 / $0.165
Kimi K2.6 $0.950 / $4.00
Kimi K2.5 $0.450 / $2.25
Llama 3.3 70B $0.100 / $0.320
MiniMax M2.7 $0.300 / $1.20

Cost = requests × (text tokens × input price + reply tokens × output price). Excludes cache and batch discounts.

How many more tokens does Thai take than English?

The same article in Thai versus English, measured on 109 stories this site publishes in both languages.

  • Qwen3.5 – 3.8 1.20×
  • Gemini 2.5 · 3.x ≈1.40×
  • Llama 4 ≈1.43×
  • Grok 4.6 · 4.7 ≈1.44×
  • Amazon Nova 2 ≈1.53×
  • DeepSeek V4 · V4.1 ≈1.57×
  • GPT-4o · 4.1 · 5.x · 6 1.67×
  • Llama 3.3 ≈1.83×
  • Mistral Medium 3.5 ≈1.96×
  • Typhoon 2.5 · Qwen3 1.99×
  • Claude Opus · Sonnet · Fable 5.x ≈2.03×
  • Claude Haiku 4.5 ≈2.73×
  • Kimi K2.5 · K2.6 ≈2.73×
  • Kimi K3 ≈2.76×
  • GLM-5.3 ≈3.03×
  • MiniMax M2.7 · M3 ≈3.47×

≈ = derived through a ratio measured against GPT (see method below)

  • Typhoon 2.5, a Thai model, uses Qwen3’s tokenizer unchanged — the file is byte-identical — so a Thai article costs 1.99× its English version — a wider gap than GPT’s 1.67×.
  • Qwen3.5 through 3.8 share a new tokenizer (248,000-entry vocabulary, up from 151,643); the Thai–English gap drops to 1.20×.
  • GPT-6 (Astra, Sol, Luna) and GPT-5.6 use the same tokenizer as GPT-5, o200k — measured through OpenRouter, the counts matched on every text — so they are counted exactly in your browser.
  • On the same Thai text, these use fewer tokens than GPT (× GPT in brackets): Qwen3.5 – 3.8 (0.77), Llama 4 (0.84), Amazon Nova 2 (0.88), Grok 4.6 · 4.7 (0.89), Gemini 2.5 · 3.x (0.92), DeepSeek V4 · V4.1 (0.94).
  • The most tokens on Thai text (× GPT): MiniMax M2.7 · M3 (2.03), Claude Opus · Sonnet · Fable 5.x (1.99), Claude Haiku 4.5 (1.87).
  • Anthropic says Claude’s new tokenizer (4.7 onward) counts about 30% more. Measured: +43% on English, only +6% on Thai.

How this was measured

  • GPT, Typhoon and Qwen are counted with the real tokenizers in your browser, tested against the reference libraries (tiktoken and Hugging Face tokenizers) on 33 real texts with zero mismatches.
  • Claude, Gemini, Grok, DeepSeek, Llama, Mistral, Kimi, GLM, MiniMax and Nova have no tokenizer we can run in your browser. We sent 12 article pairs through OpenRouter, which reports each provider’s own token count, and measured the ratio to GPT separately for Thai and English. Estimates here are GPT tokens × that ratio, with the p10–p90 spread shown as a range.
  • Models from one vendor share a row only when they measured identical on every text — Gemini 3.8 Flash and 2.5 Flash-Lite, for example. DeepSeek through OpenRouter sometimes reports exactly 79 extra tokens (a provider adds its own text); otherwise every V4 model counts the same.
  • Left out: GPT-6 Pro variants report 260–470 more input tokens than the base model, varying by text, so their cost cannot be computed from a token count. GLM-5.2 and Amazon Nova Pro counted differently from their measured siblings. Mistral Small and Cohere refused every measurement request.
  • For exact Claude counts, Anthropic’s count_tokens API is free to use (API key required).
  • Prices are pulled from OpenRouter every time the site rebuilds (daily). Latest: 2026-09-23.

FAQ

What is a token?

The unit a model splits text into — a whole word, part of a word, or a single character. API billing and context limits are counted in tokens, not characters.

Is my text sent anywhere?

No. Counting happens entirely in your browser; the tokenizer files load from this site and your text is never transmitted.

Why does Thai use more tokens?

Most tokenizers are trained mainly on English, so Thai words get split into smaller pieces — and Thai has no spaces between words, which makes splitting harder still.

How can I cut Thai token costs?

Pick a model whose tokenizer handles Thai well (Qwen3.5 – 3.8 and Llama 4 use the fewest of those measured), write long reused system prompts in English, and use prompt caching for content you send every time.