What this calculator does
Most AI pricing pages quote a price per million tokens, which is hard to turn into a budget. This calculator takes the three things you actually know about your product (how many calls you make, how big each prompt is, how long each answer is) and converts them into a monthly bill for twelve current models from Anthropic, OpenAI, Google and DeepSeek. If your app speaks its answers aloud, add the spoken characters and pick a voice provider. The speech cost is added to every row so you compare complete stacks rather than the language model alone.
How the math works
Language models bill input tokens (everything you send) and output tokens (everything the model writes) at different rates. Output is usually four to six times more expensive. For each model the calculator works out:
- Monthly requests = requests per day × 30.4 (the average month length).
- LLM cost = monthly requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000.
- Speech cost = monthly requests × characters spoken × the provider's price per million characters ÷ 1,000,000. Providers that bill per minute or per audio token are converted at about 900 characters per minute of speech.
Worked example
A support chatbot handles 2,000 conversations a day. Each call sends about 1,500 input tokens (system prompt, help-centre snippets and the customer's message) and gets back about 400 output tokens.
That is 60,800 requests a month: 91.2 million input tokens and 24.3 million output tokens. On Claude Sonnet 5.5 ($2 in / $10 out) the bill is 91.2 × $2 + 24.32 × $10 = $425.60 a month. On GPT-5.6 Luna ($0.20 / $1.20) the same traffic costs $47.42. Read 200 characters of each answer aloud with Amazon Polly Neural and you add 12.16 million characters × $16 = $194.56.
Notice how the speech line can rival the language-model line in a voice product. That is why it is worth modelling both together before you pick a stack.
Getting realistic inputs
The fastest way to estimate tokens is to paste a representative prompt into our LLM token estimator. Remember that chat apps resend the conversation history on every turn, so input tokens grow as a conversation gets longer. Reasoning models also bill their hidden "thinking" as output tokens, which can multiply the output count. For voice, our TTS cost calculator compares voice providers on their own and lets you work in minutes of audio.
Limitations
Results use standard pay-as-you-go list prices. They leave out prompt-caching discounts (often 90% off repeated input), batch discounts (usually 50%), free tiers, volume contracts, taxes and data-residency surcharges. DeepSeek rows use peak-hour rates; off-peak is half. Gemini 3.8 Flash is on a promotional price until the end of 2026. Treat the output as a planning estimate and confirm on each provider's pricing page before you commit.