AI & LLM Tools

AI API Cost Calculator

Budgeting for a language-model feature starts with simple arithmetic: how many tokens go in, how many come out, what each million costs and how many requests you expect. This calculator performs that arithmetic locally and shows input cost, output cost, the cost of a single request and the projected total for many requests. It does not include built-in model prices, so you enter the rates from your provider's pricing page and keep control of the assumptions behind the estimate.

Your workspace

Runs locally in your browser

Enter current rates from your provider's pricing page. This tool has no built-in prices, and real bills can include discounts, caching or extra fees.

Result

Your input is processed locally in your browser and is not sent to Flutters servers. Inputs are not saved by this tool.

How to use this tool

Enter input tokens, output tokens, both per-million prices and the number of requests, then choose Calculate cost. Copy the breakdown when you are done.

Example

Input

Input tokens: 2000
Output tokens: 500
Input price per 1M: 3
Output price per 1M: 15
Requests: 1000

Output

Per request: 0.0135
Input cost: 6.00
Output cost: 7.50
Total: 13.50

What does this tool do?

Most language-model APIs bill input tokens and output tokens at separate rates quoted per one million tokens. This calculator applies the formula tokens ÷ 1,000,000 × price for each side, adds them for one request, and multiplies by the number of requests. You supply the rates, so the result stays valid as providers change their pricing.

Common mistakes and limitations

The calculator contains no provider prices on purpose, because published rates change. Copy current rates from your provider's pricing page and confirm their currency and unit. Real bills can also include cached-input discounts, batch pricing, tool or search fees, minimum charges and rounding that a simple formula does not model.

The formula behind the numbers

Providers usually quote a price per one million tokens, with one rate for input and another for output. The cost of one request is input tokens divided by one million times the input rate, plus output tokens divided by one million times the output rate. Multiply by the request count for a projection. Keep the units consistent: if the rate is in dollars per million tokens, the answer is in dollars.

cost = (input_tokens / 1,000,000 × input_rate)
     + (output_tokens / 1,000,000 × output_rate)
total = cost × requests

What to measure before you estimate

Use realistic token counts: the whole prompt including system instructions and retrieved context, and the typical length of the reply rather than the maximum allowed. Output is often priced higher than input, so a verbose answer can dominate the bill. Run a few real requests and read the usage figures your provider returns, then use them here to scale up.

Related: LLM Token Counter, RAG Text Chunker

Costs this estimate does not include

Discounts for cached or batched input, fees for tools, search or image and audio tokens, minimum charges, retries and failed calls are outside this simple model. Add a safety margin, and compare the projection with your provider's billing dashboard after launch.

Frequently asked questions

Why are there no model price presets?

Prices change often and differ by model, region and plan. Entering the current rate yourself avoids showing a stale figure as if it were authoritative.

Which currency is used?

Whatever currency your price inputs use. The tool only multiplies numbers and does not convert currencies.

Where do I get token counts?

Estimate them with the LLM Token Counter, or use the usage figures returned by your provider for real requests.

Related tools