Toolantern
Home › LLM Token Estimator
Prices checked Oct 6, 2026

LLM Token Estimator

Paste a prompt, a document or a chat transcript to see roughly how many tokens it is and what it costs per call and per month on twelve models. Your text never leaves this page.

How often this text is sent or generated.
Only used when the text is input.
0estimated tokens
0words
0characters
ModelPer callPer monthPrice per 1M (in / out)

Token counts are estimates (see "How the estimate works" below). Real counts from a provider's tokenizer can differ by 10–20%, more for code or non-English text.

What is a token?

Language models do not read letters or words. They read tokens: chunks of text produced by a tokenizer, typically a common word, part of a longer word, a punctuation mark or a run of spaces. "Cat" is one token. "Unbelievably" may be three. Every provider bills by the token, and every model has a context limit measured in tokens, so knowing roughly how many tokens a prompt uses is the first step in estimating cost or checking whether a document will fit.

How the estimate works

Each provider uses its own tokenizer, and several do not publish theirs. Rather than ship megabytes of tokenizer files to your browser, this tool uses two well-established rules of thumb for English text and averages them:

Characters from Chinese, Japanese and Korean scripts are counted separately at roughly one token each, because those scripts pack far more meaning into each character. The low and high ends of the range shown come from the two rules on their own. The cost table multiplies the token estimate by each model's list price: as input, plus your expected reply length as output; or as output on its own.

Worked example

A 300-word system prompt of about 1,800 characters estimates at (1,800 ÷ 4 + 300 × 4 ÷ 3) ÷ 2 = (450 + 400) ÷ 2 = 425 tokens. Sent 10,000 times a month with a 300-token reply, that is 4.25 million input tokens and 3 million output tokens. On Claude Haiku 4.5 ($1 / $5) that costs 4.25 + 15 = $19.25 a month. On GPT-5.6 Sol ($5 / $30) it is 21.25 + 90 = $111.25.

The example also shows that output length often matters more than prompt length. Trimming replies is frequently the cheapest optimisation available. To model a whole application, with traffic, conversation history and voice output, use the AI API cost calculator.

Limitations

This is an estimator, not a tokenizer. Source code, JSON, URLs, numbers and unusual formatting usually tokenize less efficiently than prose, so expect real counts to come in higher for those. Different model families can produce noticeably different counts for the same text. For billing-grade numbers, use the token-count endpoint or usage field your provider returns with each response. Chat applications also add hidden formatting tokens around each message, and images, audio and tool definitions are billed separately and are not counted here.

Frequently asked questions

How many tokens is 1,000 words?

For typical English prose, about 1,300 to 1,400 tokens. The common rule of thumb is that one token is about three quarters of a word, or about four characters.

Why does my provider report a different token count?

Each model family has its own tokenizer, and this tool uses a fast approximation instead of any one of them. Expect differences of 10–20% on prose and more on code, data or non-English text.

Is my text uploaded anywhere?

No. Counting happens in JavaScript on this page. The text is never sent to a server, logged or stored, which makes the tool safe for confidential prompts.

Do spaces and punctuation count as tokens?

Yes. Punctuation is usually its own token, and spaces are generally merged into the token that follows them. Long runs of whitespace or line breaks can add extra tokens.

Are output tokens counted the same way as input tokens?

They are counted the same way but billed at a higher rate, typically four to six times the input price. Reasoning models also bill their internal thinking as output tokens.

How do I fit a long document into a model's context window?

Estimate its tokens here and compare with the model's context limit, leaving room for instructions and the reply. If it does not fit, split the document into chunks or retrieve only the relevant passages.

More free tools