LLM API price comparison
The same open-weight model can cost you very different amounts depending on where you call it. Below are 11 models priced in five places at once, all converted to the same unit: US dollars per 1M tokens, cache-miss, standard tier. Every figure comes from that provider's own pricing page on the date shown, and the rows where we are the expensive option are called out further down.
Method
How these numbers were put together
- One unit for everyone
- Our rates are quoted in USD credit; a $1 top-up currently becomes $20 of credit. The figures in our column are credit divided by 20 — the cash that actually leaves your card. Everyone else's figures are already cash.
- Cheapest published tier, cache-miss
- Where a provider publishes several tiers, the table takes the cheapest one that applies to a normal call: DeepSeek's off-peak rate rather than peak, DeepInfra's standard tier rather than flex. Cached input is excluded everywhere, including ours.
- Their model IDs, not ours
- Aggregators often list a dated snapshot (`-0813`) instead of the rolling model, or rename it by parameter count. Those cases are marked in the table; a dated build is not guaranteed to behave like the current one.
- Two different kinds of blank
- A cell reading "Not in catalog" means we checked and the provider does not carry that model. An empty cell means we did not verify it — not that it is unavailable. The two are never merged.
| Model ID | Y-API | Vendor official | OpenRouter | Together AI | DeepInfra |
|---|---|---|---|---|---|
| xiaomi/mimo-v2.5per 1M tokens | Y-API$0.007 Input$0.014 Output | Xiaomi MiMo$0.142a cheaper cache-hit tier also exists Input$0.28 Outputmimo-v2.5 | OpenRouter$0.1193dated snapshot build, not the rolling model Input$0.238 Outputxiaomi/mimo-v2.5-20260422 | Together AINot in catalog | DeepInfraNot in catalog |
| deepseek/deepseek-v4-flashper 1M tokens | Y-API$0.0075 Input$0.015 Output | DeepSeek$0.221,2off-peak rate; a cheaper cache-hit tier also exists Input$0.66 Outputdeepseek-v4-flash | OpenRouter$0.063dated snapshot build, not the rolling model Input$0.12 Outputdeepseek/deepseek-v4-flash-0731 | Together AI$0.143dated snapshot build, not the rolling model Input$0.28 OutputDeepSeek V4 Flash 0731 | DeepInfra$0.095standard tier; priority and flex tiers differ Input$0.18 Outputdeepseek-ai/DeepSeek-V4-Flash |
| minimax/minimax-m2.7per 1M tokens | Y-API$0.015 Input$0.06 Output | MiniMax$0.302a cheaper cache-hit tier also exists Input$1.20 OutputMiniMax-M2.7 | Together AI$0.30 Input$1.20 OutputMiniMax M2.7 | DeepInfraNot in catalog | |
| minimax/minimax-m2.5per 1M tokens | Y-API$0.015 Input$0.06 Output | MiniMax$0.302a cheaper cache-hit tier also exists Input$1.20 OutputMiniMax-M2.5 | Together AINot in catalog | DeepInfraNot in catalog | |
| qwen/qwen3.7-plusper 1M tokens | Y-API$0.0175 Input$0.065 Output | Alibaba Cloud Model StudioPrice not published | Together AI$0.32 Input$1.28 OutputQwen3.7-Plus | DeepInfraNot in catalog | |
| xiaomi/mimo-v2.5-proper 1M tokens | Y-API$0.0225 Input$0.045 Output | Xiaomi MiMo$0.4352a cheaper cache-hit tier also exists Input$0.87 Outputmimo-v2.5-pro | Together AINot in catalog | DeepInfraNot in catalog | |
| deepseek/deepseek-v4-proper 1M tokens | Y-API$0.025 Input$0.05 Output | DeepSeek$0.661,2off-peak rate; a cheaper cache-hit tier also exists Input$1.98 Outputdeepseek-v4-pro | OpenRouter$1.1223dated snapshot build, not the rolling model Input$3.366 Outputdeepseek/deepseek-v4-pro-0813 | Together AI$1.74 Input$3.48 OutputDeepSeek V4 Pro | DeepInfra$1.305standard tier; priority and flex tiers differ Input$2.60 Outputdeepseek-ai/DeepSeek-V4-Pro |
| minimax/minimax-m2.7-highspeedper 1M tokens | Y-API$0.03 Input$0.12 Output | MiniMax$0.60 Input$2.40 OutputMiniMax-M2.7-highspeed | Together AINot in catalog | DeepInfraNot in catalog | |
| minimax/minimax-m2.5-highspeedper 1M tokens | Y-API$0.03 Input$0.12 Output | MiniMax$0.60 Input$2.40 OutputMiniMax-M2.5-highspeed | Together AINot in catalog | DeepInfraNot in catalog | |
| qwen/qwen3.7-maxper 1M tokens | Y-API$0.075 Input$0.225 Output | Alibaba Cloud Model StudioPrice not published | Together AI$1.25 Input$3.75 OutputQwen3.7-Max | DeepInfra$2.505standard tier; priority and flex tiers differ Input$7.50 OutputQwen/Qwen3.7-Max | |
| qwen/qwen3.8-maxper 1M tokens | Y-API$0.10 Input$0.30 Output | Alibaba Cloud Model StudioPrice not published | OpenRouter$2.00 Input$6.00 Output | Together AI$2.504listed under a different name Input$6.25 OutputQwen3.8-2.4T-A95B | DeepInfra$1.655standard tier; priority and flex tiers differ Input$4.951 OutputQwen/Qwen3.8-Max |
- 1off-peak rate
- 2a cheaper cache-hit tier also exists
- 3dated snapshot build, not the rolling model
- 4listed under a different name
- 5standard tier; priority and flex tiers differ
Where we are cheaper
The gap comes from the top-up rate, not from marked-down rates
Read our credit rates as if they were dollars and they sit roughly on top of the vendors' own list prices — identical for all four MiniMax models and for xiaomi/mimo-v2.5, a little above official for xiaomi/mimo-v2.5-pro, below it for the two DeepSeek builds. Practically the whole difference is the 20:1 top-up rate. That is also why it is honest to say the multiple is temporary: at 1:10 it halves.
Against vendor official prices, cash for cash, the models we can compare land between 19.3× and 29.3× cheaper on input.
Where we cost more
Three cases where you should not use us
Workloads that hit the prompt cache constantly
DeepSeek bills cached input at $0.022 per 1M off-peak and Xiaomi at $0.0028 — below our cash rate for the same models. We have no cache-hit tier at all. If you replay a long, stable system prompt thousands of times, calling the vendor directly can cost less than calling us.
Models outside our catalog
Our catalog is 16 model IDs. GPT, Claude and Gemini are not in it, and neither is most of the long tail an aggregator carries. No price comparison helps if the model you need is not on the list.
When you need a price that cannot change
The multiple in this table rests on a limited-time 20:1 top-up rate. Credit already in your account is not repriced, but new top-ups after the promotion ends convert at 1:10, which halves every advantage shown here. Vendor list prices are also not contractual, but they move slowly and in public.
Gaps
What is missing from this table, and why
- Not every model is here
- 5 of our 16 models are absent. Two Qwen models have no verified external price of any kind, two DeepSeek builds appear on no pricing page we could check, and one model does not disclose its vendor.
- Alibaba publishes no Qwen price
- Its public model pages list capabilities only; per-token billing sits behind a region-specific console. That is why the official column is blank on every Qwen row — a blank is better than a number copied from a blog. Our own rates are on a public page and in a machine-readable endpoint.
- No third-party figures were used
- For one DeepSeek model alone, three widely-cited posts quote three mutually incompatible prices. Only a provider's own pricing page counts as a source here.
- Aggregators disagree by up to 2×
- Qwen3.7-Max is $1.25 per 1M input on Together AI and $2.50 on DeepInfra — same model, same unit, same day. Treat any single aggregator figure as one data point, not as the market price.
FAQ
Questions this table raises
- Which is the cheapest way to call DeepSeek V4 Pro?
- Verified 2026-08-26, per 1M input tokens: Y-API $0.025 in cash, DeepSeek's own API $0.66 off-peak and $1.32 peak, DeepInfra $1.30, OpenRouter $1.122 for the dated 0813 build, Together AI $1.74. The one case that flips is a constantly cached prompt — DeepSeek charges $0.022 per 1M for cached input, slightly below our rate.
- Why are Y-API's prices quoted in credit instead of dollars?
- Credit is the metering unit inside the service: a top-up converts dollars into credit at the current rate, and each model deducts credit per token. Quoting the credit rate keeps the per-model figure stable when the top-up rate changes. To get the cash price, divide the credit rate by the current rate — 20 today, 10 after the promotion. Both figures are published, and the cash price is in the machine-readable pricing endpoint.
- Is an aggregator ever cheaper than the model vendor itself?
- Yes, and this table has an example. OpenRouter lists the dated xiaomi/mimo-v2.5-20260422 build at $0.119 per 1M input, below Xiaomi's own $0.14. Aggregators also sometimes list only a dated snapshot, which will not track the rolling model as it is updated — cheaper, but not the same thing.
- How much of the difference is a real discount versus the top-up promotion?
- Almost all of it is the promotion. Compared as if credit were dollars, our rates equal the vendors' own list prices for all four MiniMax models and for xiaomi/mimo-v2.5, sit 3% above official for xiaomi/mimo-v2.5-pro, and below it for the two DeepSeek builds. The large multiples in the table come from the 20:1 top-up rate, which is limited-time and will revert to 1:10.
- How often is this table updated?
- Every figure carries the date it was checked, shown at the top of the page. Providers change prices without notice, so treat anything older than a few weeks as a starting point and follow the source link before committing spend. Our own rates are synced from the live catalog at build time, so they are never staler than the last deploy.
Provenance
Sources
One link per provider, each pointing at the page the figures were read from. Y-API's own rates come from its models and pricing pages, and from https://y-api.bestvirtualgoods.com/pricing.json.
Check one model against your own bill
Pick the model you call most, find it in the table, and multiply by last month's token count. If the arithmetic works out, sign in and point one call here — the request log will show the real deduction.