Monthly cost statement
How much do LLM APIs actually cost per month?
Token prices lie. The same support reply can cost 17× more on one model than another — and cache hit rates change everything. Pick your workload, pick your models, get the real number.
Usage
Workload preset
Inputs
Cache math: cached input tokens are billed at 90% off — an illustrative assumption; real cache discounts vary by provider.
Models
Amount due
$0—
per month, for this workload
You save
$0vs —
per month, at this workload
Itemized charges
| Model | / month | vs cheapest |
|---|
Relative scale
Input vs output rates
Log–log. Position is the list price per 1M tokens; hover a dot for this workload’s monthly cost.
Cache leverage
Monthly cost as the cache hit rate moves 0–90%. The rule marks your current setting.
Statement guide
How to read your statement
-
Set your workload
Pick a preset, or type your own numbers. Requests / month is how many API calls you make. Input tokens / request is what you send — prompts and retrieved context. Output tokens / request is what the model writes back. Cache hit rate is the share of input tokens served from cache, billed at 90% off.
-
Choose your models
Toggle the model chips to build your comparison set. The dot on each chip is the provider’s color — it is the same color that model wears in every chart below, so you can track it across the page.
-
Read the bill
Amount due is the cheapest model for your workload. You save is what you keep by not picking the most expensive one. Each itemized row shows its vs cheapest multiple — a 17× means you pay seventeen times more for the same workload. The charts show the same numbers at different scales: relative size, input/output price geometry, and how cache changes the ranking.
Master rates
Rate card
USD per 1M tokens. Updated 2026-10-03. Prices change frequently — verify on the provider's pricing page before production use.
| Model | Provider | Input $/1M | Output $/1M | Context |
|---|