# LLM API price comparison — Y-API

> The same models priced side by side across vendor official APIs, OpenRouter, Together AI, DeepInfra and Y-API — dollars per 1M tokens, every figure dated and sourced.

This is the markdown representation of https://y-api.bestvirtualgoods.com/compare/prices. All external prices verified 2026-08-26 against the pricing pages linked at the bottom. Providers change prices without notice — check the source before making a decision on this table alone. Generated by `scripts/generate-seo-assets.mjs` from the same copy the page renders — do not edit by hand.

## Summary

The same model can cost you very different amounts depending on where you call it. Below are 16 models priced in five places at once, all converted to the same unit: US dollars per 1M tokens, cache-miss, standard tier. Every figure comes from that provider's own pricing page on the date shown, and the rows where we are the expensive option are called out further down.

## The table

Every figure is USD / 1M tokens, input / output. Covers all 16 paid models.

| Model ID | Y-API (USD / 1M tokens) | Vendor official | OpenRouter | Together AI | DeepInfra |
| --- | --- | --- | --- | --- | --- |
| `deepseek/deepseek-v4-pro` | $0.05 / $0.10 | $0.66 / $1.98[1][2] | $0.6587 / $1.976[2][3] | $1.32 / $3.96[2][3] | $1.30 / $2.60[5] |
| `deepseek/deepseek-v4.1-flash` | $0.02 / $0.10 | $0.15 / $0.60[1][2] | $0.15 / $0.60[2] |  |  |
| `deepseek/deepseek-v4-flash-0731` | $0.015 / $0.03 | $0.22 / $0.66[1][2][4] |  |  |  |
| `qwen/qwen3.8-flash` | $0.02 / $0.05 | $0.113 / $0.382[6] | $0.15 / $0.47[2] | $0.15 / $0.47 | Not in catalog |
| `z-ai/glm-5.2` | $0.14 / $0.44 | $1.40 / $4.40[2] |  |  |  |
| `z-ai/glm-5.3` | $0.14 / $0.50 | $1.40 / $4.40[2] |  |  |  |
| `z-ai/glm-5.3-flash` | $0.015 / $0.05 | $0.15 / $0.50[2] |  |  |  |
| `moonshotai/kimi-k3` | $0.30 / $1.50 | $3.00 / $15.00[2] |  |  |  |
| `moonshotai/kimi-k2.6` | $0.095 / $0.40 | $0.95 / $4.00[2] |  |  |  |
| `anthropic/claude-opus-5` | $0.50 / $2.50 | $5.00 / $25.00[2] |  |  |  |
| `anthropic/claude-sonnet-5` | $0.20 / $1.00 | $2.00 / $10.00[2] |  |  |  |
| `openai/gpt-5.6-sol` | $0.50 / $3.00 | $4.00 / $20.00[2][5][7][8] |  |  |  |
| `openai/gpt-5.6-luna` | $0.03 / $0.13 | $0.20 / $1.20[2][5][7] |  |  |  |
| `tencent/hy4-preview` | $0.10 / $0.30 | $1.20 / $3.60 |  |  |  |
| `stepfun/step-3.7-flash` | $0.02 / $0.12 | $0.22 / $1.32 |  |  |  |
| `xiaomi/mimo-v2.6-flash` | $0.018 / $0.036 | $0.16 / $0.32 |  |  |  |

Caveats referenced above:

1. off-peak rate
2. a cheaper cache-hit tier also exists
3. dated snapshot build, not the rolling model
4. historical official price for this version; current rolling-model pricing may differ
5. standard tier; priority and flex tiers differ
6. cheapest of its six regional price tables
7. cheapest input-length tier; a longer prompt costs more per token
8. limited-time promotional rate; the list price is higher

## The gap comes from the top-up rate, not from marked-down rates

Read our credit rates as if they were dollars and they sit roughly on top of the vendors' own list prices: across the models we can compare, ours come to between 0.682× and 1.77× the vendor's own figure. Practically the whole difference is the 1:10 top-up rate.

## Two cases where you should not use us

### Workloads that hit the prompt cache constantly

DeepSeek bills cached input on deepseek/deepseek-v4.1-flash at $0.003 per 1M against our $0.02, DeepSeek bills cached input on deepseek/deepseek-v4-pro at $0.022 per 1M against our $0.05, OpenAI bills cached input on openai/gpt-5.6-luna at $0.02 per 1M against our $0.03 and OpenAI bills cached input on openai/gpt-5.6-sol at $0.40 per 1M against our $0.50. We have no cache-hit tier at all. If you replay a long, stable system prompt thousands of times, calling the vendor directly can cost less than calling us.

### Models outside our catalog

Our catalog is 20 model IDs. Gemini is not in it, and neither is most of the long tail an aggregator carries. No price comparison helps if the model you need is not on the list.

## How these numbers were put together

- **One unit for everyone**: Our rates are quoted in USD credit; a $1 top-up becomes $10 of credit. The figures in our column are credit divided by 10 — the cash that actually leaves your card. Everyone else's figures are already cash.
- **Cheapest published tier, cache-miss**: Where a provider publishes several tiers, the table takes the cheapest one that applies to a normal call: DeepSeek's off-peak rate rather than peak, DeepInfra's standard tier rather than flex. Cached input is excluded everywhere, including ours.
- **Their model IDs, not ours**: Aggregators often list a dated snapshot (`-0813`) instead of the rolling model, or rename it by parameter count. Those cases are marked in the table; a dated build is not guaranteed to behave like the current one.
- **Two different kinds of blank**: A cell reading "Not in catalog" means we checked and the provider does not carry that model. An empty cell means we did not verify it — not that it is unavailable. The two are never merged.

## Coverage and pricing limits

- **Model coverage**: The table covers all 16 paid models. Historical version prices and billing tiers are marked beside the figures.
- **Alibaba does not have one Qwen price**: Model Studio splits the same build across six regional rate cards, then across input-length tiers, then discounts some builds for a limited time and others at night, and bills cache hits separately at a rate the page never states. The official column takes the cheapest cell that applies to an ordinary real-time call — cheapest region, shortest input tier, after the limited-time discount, cache-miss — which is the reading least favourable to us. The night tier is left out on purpose: the English page says "night 80% off" while the Chinese wording reads as paying 80% at night, and we will not pick whichever reading flatters us.
- **No third-party figures were used**: For one DeepSeek model alone, three widely-cited posts quote three mutually incompatible prices. Only a provider's own pricing page counts as a source here.
- **Aggregators disagree by up to 2×**: The deepseek/deepseek-v4-pro row reads $0.6587 per 1M input on OpenRouter and $1.32 on Together AI — one model, one unit, both read from that provider’s own page. Treat any single aggregator figure as one data point, not as the market price.

## Compared against a single provider

This table is wide and shallow on purpose: one figure per cell, no room to explain how any one provider structures its rates. Each page below takes a single provider and goes the other way — every model both sides sell, the tiers sitting behind each figure, the date it was read, and the cases where that provider is the cheaper of the two.

- [OpenRouter alternative: Y-API vs OpenRouter](https://y-api.bestvirtualgoods.com/vs/openrouter)
- [Y-API vs the DeepSeek API](https://y-api.bestvirtualgoods.com/vs/deepseek-official)
- [Y-API vs the official Qwen API](https://y-api.bestvirtualgoods.com/vs/qwen-official)
- [Y-API vs Xiaomi MiMo](https://y-api.bestvirtualgoods.com/vs/xiaomi-official)
- [Y-API vs Together AI](https://y-api.bestvirtualgoods.com/vs/together)
- [Y-API vs DeepInfra](https://y-api.bestvirtualgoods.com/vs/deepinfra)

## Questions this table raises

### What is the cheapest way to call deepseek/deepseek-v4-pro?

Per 1M input tokens, every figure read on or after 2026-08-26: Y-API $0.05 in cash, DeepSeek's own API $0.66, OpenRouter $0.6587 listed as deepseek/deepseek-v4-pro-0813, Together AI $1.32 listed as DeepSeek V4 Pro 0813 and DeepInfra $1.30 listed as deepseek-ai/DeepSeek-V4-Pro. The one case that flips is a constantly cached prompt — DeepSeek charges $0.022 per 1M for cached input, below our $0.05.

### Why are Y-API's prices quoted in credit instead of dollars?

Credit is the metering unit inside the service: a top-up converts dollars into credit at the current rate, and each model deducts credit per token. Quoting the credit rate keeps the per-model figure stable when the top-up rate changes. To get the cash price, divide the credit rate by the top-up rate (10). Both figures are published, and the cash price is in the machine-readable pricing endpoint.

### Is an aggregator ever cheaper than the model vendor itself?

Yes, and this table has an example. OpenRouter lists deepseek/deepseek-v4-pro-0813 at $0.6587 per 1M input, below DeepSeek's own $0.66. Aggregators also sometimes list only a dated snapshot, which will not track the rolling model as it is updated — cheaper, but not the same thing.

### How much of the difference is a real discount versus the top-up rate?

Almost all of it is the top-up rate. Read as if credit were dollars, our rates come to between 0.682× and 1.77× the vendors' own list prices across the models we can compare. The large multiples in the table come from the 1:10 top-up rate.

### How often is this table updated?

Every figure carries the date it was checked, shown at the top of the page. Providers change prices without notice, so treat anything older than a few weeks as a starting point and follow the source link before committing spend. Our own rates are synced from the live catalog at build time, so they are never staler than the last deploy.

## Sources

- [DeepSeek](https://api-docs.deepseek.com/quick_start/pricing/)
- [Z.ai](https://docs.z.ai/guides/overview/pricing)
- [Kimi](https://platform.kimi.ai/docs/pricing/chat-k3)
- [OpenAI](https://developers.openai.com/api/docs/pricing)
- [Anthropic](https://www.anthropic.com/pricing#api)
- [MiniMax](https://platform.minimax.io/docs/guides/pricing-paygo)
- [Xiaomi MiMo](https://mimo.mi.com/docs/en-US/price/pay-as-you-go)
- [Tencent Hunyuan](https://cloud.tencent.com/document/product/1729/97731)
- [StepFun](https://www.stepfun.com/pricing)
- [Alibaba Cloud Model Studio](https://www.alibabacloud.com/help/en/model-studio/billing-for-model-studio)
- [OpenRouter](https://openrouter.ai/models)
- [Together AI](https://www.together.ai/pricing)
- [DeepInfra](https://deepinfra.com/pricing)
- [DeepSeek · deepseek-v4-flash](https://api-docs.deepseek.com/news/news260813/)

## Links

- HTML version of this page: https://y-api.bestvirtualgoods.com/compare/prices
- Site index for agents: https://y-api.bestvirtualgoods.com/llms.txt
- Full reference (single file): https://y-api.bestvirtualgoods.com/llms-full.txt
- OpenAPI 3.1 spec: https://y-api.bestvirtualgoods.com/openapi.json
- Model catalog (JSON, no key needed): https://y-api.bestvirtualgoods.com/models.json
- API base URL: `https://api.y-api.bestvirtualgoods.com/v1`
- Contact: support@bestvirtualgoods.com
- Cost calculator for your own volume: https://y-api.bestvirtualgoods.com/calculator
- Structured prices (JSON, ours and vendor official): https://y-api.bestvirtualgoods.com/pricing.json
- OpenRouter price list (verified 2026-09-18): https://openrouter.ai/models
- Together AI price list (verified 2026-08-26): https://www.together.ai/pricing
- DeepInfra price list (verified 2026-08-26): https://deepinfra.com/pricing
