Toolantern
Home › AI API Cost Calculator
Prices checked Oct 6, 2026

AI API Cost Calculator

Describe your workload once and see what it costs per month on every major model, including voice output if your app talks back. Everything runs in your browser.

Traffic
One request = one model call.
Tokens per request
Prompt + history + retrieved context.
The model's reply (incl. reasoning tokens).
Speech output (optional)
0 = no text-to-speech. ~900 characters ≈ 1 minute.
—requests per month
—cheapest model / month
—most expensive / month
ModelLLM / moSpeech / moTotal / moPer 1k requests

What this calculator does

Most AI pricing pages quote a price per million tokens, which is hard to turn into a budget. This calculator takes the three things you actually know about your product (how many calls you make, how big each prompt is, how long each answer is) and converts them into a monthly bill for twelve current models from Anthropic, OpenAI, Google and DeepSeek. If your app speaks its answers aloud, add the spoken characters and pick a voice provider. The speech cost is added to every row so you compare complete stacks rather than the language model alone.

How the math works

Language models bill input tokens (everything you send) and output tokens (everything the model writes) at different rates. Output is usually four to six times more expensive. For each model the calculator works out:

Worked example

A support chatbot handles 2,000 conversations a day. Each call sends about 1,500 input tokens (system prompt, help-centre snippets and the customer's message) and gets back about 400 output tokens.

That is 60,800 requests a month: 91.2 million input tokens and 24.3 million output tokens. On Claude Sonnet 5.5 ($2 in / $10 out) the bill is 91.2 × $2 + 24.32 × $10 = $425.60 a month. On GPT-5.6 Luna ($0.20 / $1.20) the same traffic costs $47.42. Read 200 characters of each answer aloud with Amazon Polly Neural and you add 12.16 million characters × $16 = $194.56.

Notice how the speech line can rival the language-model line in a voice product. That is why it is worth modelling both together before you pick a stack.

Getting realistic inputs

The fastest way to estimate tokens is to paste a representative prompt into our LLM token estimator. Remember that chat apps resend the conversation history on every turn, so input tokens grow as a conversation gets longer. Reasoning models also bill their hidden "thinking" as output tokens, which can multiply the output count. For voice, our TTS cost calculator compares voice providers on their own and lets you work in minutes of audio.

Limitations

Results use standard pay-as-you-go list prices. They leave out prompt-caching discounts (often 90% off repeated input), batch discounts (usually 50%), free tiers, volume contracts, taxes and data-residency surcharges. DeepSeek rows use peak-hour rates; off-peak is half. Gemini 3.8 Flash is on a promotional price until the end of 2026. Treat the output as a planning estimate and confirm on each provider's pricing page before you commit.

Frequently asked questions

How accurate is this AI API cost estimate?

The arithmetic is exact for the prices listed. The uncertainty comes from your inputs: real token counts vary from request to request, so use an average from your logs if you have them. Prices were checked against each provider's official pricing page on the date shown at the top of the page.

Why is output more expensive than input?

Generating each output token requires a full forward pass of the model, one token at a time, while input tokens are processed in parallel. Providers price output at roughly four to six times the input rate to reflect that.

Does the calculator include prompt caching or batch discounts?

No. It uses standard list prices so models are compared on equal terms. If you resend a large, unchanging system prompt, caching can cut input cost sharply, so treat the input share of the bill as an upper bound.

How many characters is one minute of speech?

Conversational English runs at about 150 words per minute, and an average word plus its space is about six characters, so roughly 900 characters make one minute of audio. Fast or slow voices will shift this.

Which model is cheapest for a chatbot?

For pure cost, budget-tier models such as GPT-5.6 Luna, DeepSeek V4.1 Flash and Gemini 3.5 Flash-Lite are usually cheapest. Whether they are good enough depends on your task, so test quality on a sample of real conversations before switching.

Is my data sent anywhere?

No. The calculator is plain JavaScript running in your browser. Nothing you enter is uploaded or stored.

More free tools