LLM API price comparison
The same model can cost you very different amounts depending on where you call it. Below are 16 models priced in five places at once, all converted to the same unit: US dollars per 1M tokens, cache-miss, standard tier. Every figure comes from that provider's own pricing page on the date shown, and the rows where we are the expensive option are called out further down.
Method
How these numbers were put together
- One unit for everyone
- Our rates are quoted in USD credit; a $1 top-up becomes $10 of credit. The figures in our column are credit divided by 10 — the cash that actually leaves your card. Everyone else's figures are already cash.
- Cheapest published tier, cache-miss
- Where a provider publishes several tiers, the table takes the cheapest one that applies to a normal call: DeepSeek's off-peak rate rather than peak, DeepInfra's standard tier rather than flex. Cached input is excluded everywhere, including ours.
- Their model IDs, not ours
- Aggregators often list a dated snapshot (`-0813`) instead of the rolling model, or rename it by parameter count. Those cases are marked in the table; a dated build is not guaranteed to behave like the current one.
- Two different kinds of blank
- A cell reading "Not in catalog" means we checked and the provider does not carry that model. An empty cell means we did not verify it — not that it is unavailable. The two are never merged.
| Model ID | Y-API | Vendor official | OpenRouter | Together AI | DeepInfra |
|---|---|---|---|---|---|
| Vendor: DeepSeek | |||||
| deepseek/deepseek-v4-proper 1M tokens | Y-API$0.05 Input$0.10 Output | DeepSeek$0.661,2off-peak rate; a cheaper cache-hit tier also exists Input$1.98 Outputdeepseek-v4-pro | OpenRouter$0.65873,2dated snapshot build, not the rolling model; a cheaper cache-hit tier also exists Input$1.976 Outputdeepseek/deepseek-v4-pro-0813 | Together AI$1.323,2dated snapshot build, not the rolling model; a cheaper cache-hit tier also exists Input$3.96 OutputDeepSeek V4 Pro 0813 | DeepInfra$1.305standard tier; priority and flex tiers differ Input$2.60 Outputdeepseek-ai/DeepSeek-V4-Pro |
| deepseek/deepseek-v4.1-flashper 1M tokens | Y-API$0.02 Input$0.10 Output | DeepSeek$0.151,2off-peak rate; a cheaper cache-hit tier also exists Input$0.60 Outputdeepseek-flash | OpenRouter$0.152a cheaper cache-hit tier also exists Input$0.60 Output | ||
| deepseek/deepseek-v4-flash-0731per 1M tokens | Y-API$0.015 Input$0.03 Output | DeepSeek$0.221,2,4off-peak rate; a cheaper cache-hit tier also exists; historical official price for this version; current rolling-model pricing may differ Input$0.66 Outputdeepseek-v4-flash | |||
| Vendor: Qwen | |||||
| qwen/qwen3.8-flashper 1M tokens | Y-API$0.02 Input$0.05 Output | Alibaba Cloud Model Studio$0.1136cheapest of its six regional price tables Input$0.382 Outputqwen3.8-flash | OpenRouter$0.152a cheaper cache-hit tier also exists Input$0.47 Output | Together AI$0.15 Input$0.47 OutputQwen3.8 Flash | DeepInfraNot in catalog |
| Vendor: Z.ai | |||||
| z-ai/glm-5.2per 1M tokens | Y-API$0.14 Input$0.44 Output | Z.ai$1.402a cheaper cache-hit tier also exists Input$4.40 OutputGLM-5.2 | |||
| z-ai/glm-5.3per 1M tokens | Y-API$0.14 Input$0.50 Output | Z.ai$1.402a cheaper cache-hit tier also exists Input$4.40 OutputGLM-5.3 | |||
| z-ai/glm-5.3-flashper 1M tokens | Y-API$0.015 Input$0.05 Output | Z.ai$0.152a cheaper cache-hit tier also exists Input$0.50 OutputGLM-5.3-Flash | |||
| Vendor: Moonshot AI | |||||
| moonshotai/kimi-k3per 1M tokens | Y-API$0.30 Input$1.50 Output | Kimi$3.002a cheaper cache-hit tier also exists Input$15.00 Outputkimi-k3 | |||
| moonshotai/kimi-k2.6per 1M tokens | Y-API$0.095 Input$0.40 Output | Kimi$0.952a cheaper cache-hit tier also exists Input$4.00 Outputkimi-k2.6 | |||
| Vendor: Anthropic | |||||
| anthropic/claude-opus-5per 1M tokens | Y-API$0.50 Input$2.50 Output | Anthropic$5.002a cheaper cache-hit tier also exists Input$25.00 Outputclaude-opus-5 | |||
| anthropic/claude-sonnet-5per 1M tokens | Y-API$0.20 Input$1.00 Output | Anthropic$2.002a cheaper cache-hit tier also exists Input$10.00 Outputclaude-sonnet-5 | |||
| Vendor: OpenAI | |||||
| openai/gpt-5.6-solper 1M tokens | Y-API$0.50 Input$3.00 Output | OpenAI$4.002,5,7,8a cheaper cache-hit tier also exists; standard tier; priority and flex tiers differ; cheapest input-length tier; a longer prompt costs more per token; limited-time promotional rate; the list price is higher Input$20.00 Outputgpt-5.6-sol | |||
| openai/gpt-5.6-lunaper 1M tokens | Y-API$0.03 Input$0.13 Output | OpenAI$0.202,5,7a cheaper cache-hit tier also exists; standard tier; priority and flex tiers differ; cheapest input-length tier; a longer prompt costs more per token Input$1.20 Outputgpt-5.6-luna | |||
| Vendor: Other | |||||
| tencent/hy4-previewper 1M tokens | Y-API$0.10 Input$0.30 Output | Tencent Hunyuan$1.20 Input$3.60 Outputhy4-preview | |||
| stepfun/step-3.7-flashper 1M tokens | Y-API$0.02 Input$0.12 Output | StepFun$0.22 Input$1.32 Outputstep-3.7-flash | |||
| xiaomi/mimo-v2.6-flashper 1M tokens | Y-API$0.018 Input$0.036 Output | Xiaomi MiMo$0.16 Input$0.32 Outputmimo-v2.6-flash | |||
- 1off-peak rate
- 2a cheaper cache-hit tier also exists
- 3dated snapshot build, not the rolling model
- 4historical official price for this version; current rolling-model pricing may differ
- 5standard tier; priority and flex tiers differ
- 6cheapest of its six regional price tables
- 7cheapest input-length tier; a longer prompt costs more per token
- 8limited-time promotional rate; the list price is higher
Where we are cheaper
The gap comes from the top-up rate, not from marked-down rates
Read our credit rates as if they were dollars and they sit roughly on top of the vendors' own list prices: across the models we can compare, ours come to between 0.682× and 1.77× the vendor's own figure. Practically the whole difference is the 1:10 top-up rate.
Against vendor official prices, cash for cash, the models we can compare land between 5.65× and 14.7× cheaper on input.
Where we cost more
Two cases where you should not use us
Workloads that hit the prompt cache constantly
DeepSeek bills cached input on deepseek/deepseek-v4.1-flash at $0.003 per 1M against our $0.02, DeepSeek bills cached input on deepseek/deepseek-v4-pro at $0.022 per 1M against our $0.05, OpenAI bills cached input on openai/gpt-5.6-luna at $0.02 per 1M against our $0.03 and OpenAI bills cached input on openai/gpt-5.6-sol at $0.40 per 1M against our $0.50. We have no cache-hit tier at all. If you replay a long, stable system prompt thousands of times, calling the vendor directly can cost less than calling us.
Models outside our catalog
Our catalog is 20 model IDs. Gemini is not in it, and neither is most of the long tail an aggregator carries. No price comparison helps if the model you need is not on the list.
Gaps
Coverage and pricing limits
- Model coverage
- The table covers all 16 paid models. Historical version prices and billing tiers are marked beside the figures.
- Alibaba does not have one Qwen price
- Model Studio splits the same build across six regional rate cards, then across input-length tiers, then discounts some builds for a limited time and others at night, and bills cache hits separately at a rate the page never states. The official column takes the cheapest cell that applies to an ordinary real-time call — cheapest region, shortest input tier, after the limited-time discount, cache-miss — which is the reading least favourable to us. The night tier is left out on purpose: the English page says "night 80% off" while the Chinese wording reads as paying 80% at night, and we will not pick whichever reading flatters us.
- No third-party figures were used
- For one DeepSeek model alone, three widely-cited posts quote three mutually incompatible prices. Only a provider's own pricing page counts as a source here.
- Aggregators disagree by up to 2×
- The deepseek/deepseek-v4-pro row reads $0.6587 per 1M input on OpenRouter and $1.32 on Together AI — one model, one unit, both read from that provider’s own page. Treat any single aggregator figure as one data point, not as the market price.
One provider at a time
Compared against a single provider
This table is wide and shallow on purpose: one figure per cell, no room to explain how any one provider structures its rates. Each page below takes a single provider and goes the other way — every model both sides sell, the tiers sitting behind each figure, the date it was read, and the cases where that provider is the cheaper of the two.
FAQ
Questions this table raises
- What is the cheapest way to call deepseek/deepseek-v4-pro?
- Per 1M input tokens, every figure read on or after 2026-08-26: Y-API $0.05 in cash, DeepSeek's own API $0.66, OpenRouter $0.6587 listed as deepseek/deepseek-v4-pro-0813, Together AI $1.32 listed as DeepSeek V4 Pro 0813 and DeepInfra $1.30 listed as deepseek-ai/DeepSeek-V4-Pro. The one case that flips is a constantly cached prompt — DeepSeek charges $0.022 per 1M for cached input, below our $0.05.
- Why are Y-API's prices quoted in credit instead of dollars?
- Credit is the metering unit inside the service: a top-up converts dollars into credit at the current rate, and each model deducts credit per token. Quoting the credit rate keeps the per-model figure stable when the top-up rate changes. To get the cash price, divide the credit rate by the top-up rate (10). Both figures are published, and the cash price is in the machine-readable pricing endpoint.
- Is an aggregator ever cheaper than the model vendor itself?
- Yes, and this table has an example. OpenRouter lists deepseek/deepseek-v4-pro-0813 at $0.6587 per 1M input, below DeepSeek's own $0.66. Aggregators also sometimes list only a dated snapshot, which will not track the rolling model as it is updated — cheaper, but not the same thing.
- How much of the difference is a real discount versus the top-up rate?
- Almost all of it is the top-up rate. Read as if credit were dollars, our rates come to between 0.682× and 1.77× the vendors' own list prices across the models we can compare. The large multiples in the table come from the 1:10 top-up rate.
- How often is this table updated?
- Every figure carries the date it was checked, shown at the top of the page. Providers change prices without notice, so treat anything older than a few weeks as a starting point and follow the source link before committing spend. Our own rates are synced from the live catalog at build time, so they are never staler than the last deploy.
Provenance
Sources
Sources include provider pricing pages and version-specific pricing announcements. Our own prices are published on the models page, pricing page, and in pricing.json.
Check one model against your own bill
Pick the model you call most, find it in the table, and multiply by last month's token count. If the arithmetic works out, sign in and point one call here — the request log will show the real deduction.