What is a token?
Language models do not read letters or words. They read tokens: chunks of text produced by a tokenizer, typically a common word, part of a longer word, a punctuation mark or a run of spaces. "Cat" is one token. "Unbelievably" may be three. Every provider bills by the token, and every model has a context limit measured in tokens, so knowing roughly how many tokens a prompt uses is the first step in estimating cost or checking whether a document will fit.
How the estimate works
Each provider uses its own tokenizer, and several do not publish theirs. Rather than ship megabytes of tokenizer files to your browser, this tool uses two well-established rules of thumb for English text and averages them:
- About 4 characters per token. Characters ÷ 4.
- About ¾ of a word per token. Words × 4 ÷ 3.
Characters from Chinese, Japanese and Korean scripts are counted separately at roughly one token each, because those scripts pack far more meaning into each character. The low and high ends of the range shown come from the two rules on their own. The cost table multiplies the token estimate by each model's list price: as input, plus your expected reply length as output; or as output on its own.
Worked example
A 300-word system prompt of about 1,800 characters estimates at (1,800 ÷ 4 + 300 × 4 ÷ 3) ÷ 2 = (450 + 400) ÷ 2 = 425 tokens. Sent 10,000 times a month with a 300-token reply, that is 4.25 million input tokens and 3 million output tokens. On Claude Haiku 4.5 ($1 / $5) that costs 4.25 + 15 = $19.25 a month. On GPT-5.6 Sol ($5 / $30) it is 21.25 + 90 = $111.25.
The example also shows that output length often matters more than prompt length. Trimming replies is frequently the cheapest optimisation available. To model a whole application, with traffic, conversation history and voice output, use the AI API cost calculator.
Limitations
This is an estimator, not a tokenizer. Source code, JSON, URLs, numbers and unusual formatting usually tokenize less efficiently than prose, so expect real counts to come in higher for those. Different model families can produce noticeably different counts for the same text. For billing-grade numbers, use the token-count endpoint or usage field your provider returns with each response. Chat applications also add hidden formatting tokens around each message, and images, audio and tool definitions are billed separately and are not counted here.