OpenRating is an independent, open-source AI rating platform. Join the community

OOpenRating

Tools

Token Calculator

Count tokens with real GPT-family tokenizers, then see what any request costs on every model OpenRating rates.

How token counting works

This calculator runs the same byte-pair encoding tokenizers used by OpenAI models — o200k_base for GPT-5 and GPT-4o, cl100k_base for GPT-4 and GPT-3.5 — entirely in your browser. Your text never leaves the page. Claude, Gemini, and DeepSeek use their own vocabularies, so treat counts for those models as estimates within roughly 5–10%.

Reading the estimates

The cost table multiplies your token counts by each model's published input and output prices per million tokens. It does not include prompt-caching discounts, which can cut input costs substantially on providers like Anthropic and OpenAI. For end-to-end economics — including chain-of-thought tokens and cache behavior — see the task cost measurements on the leaderboard.

Frequently asked questions

What is a token?

A token is the unit LLMs use to read text — roughly a word fragment. 'OpenRating' is about two tokens, and a typical English sentence of 20 words is around 30 tokens. Providers bill APIs per token (usually per million tokens), so accurate token counts matter for cost estimation.

Why do token counts differ between models?

Every model family uses its own tokenizer vocabulary. GPT-5 and GPT-4o use the o200k_base encoding, older GPT-4 and GPT-3.5 models use cl100k_base, and Claude, Gemini, and DeepSeek each use their own vocabularies. Counts across providers typically differ by 5–10%. This calculator uses GPT-family encodings for all models, so treat non-OpenAI counts as approximations.

How is the estimated cost calculated?

For each model we multiply your input tokens by its input price and your estimated response tokens by its output price, both priced per million tokens. Actual costs can be lower when prompt caching applies — cached input tokens are billed at a fraction of the standard rate on most providers.

How can I reduce token costs?

Shorten prompts, reuse system prompts so providers can cache them (cache hits bill at a fraction of standard input rates), pick cheaper models for high-volume tasks, and cap output length. For models with heavy chain-of-thought, the task cost column on the leaderboard shows real per-run economics beyond raw token prices.

Benchmark Newsletter

New benchmark ratings in your inbox

Monthly empirical updates when labs drop new models and prices change. No marketing spam.

Join 4,200+ ML engineers and researchers. Unsubscribe at any time.