# Y-API — full reference > OpenAI-compatible LLM API for 20+ models. Pay per token. Compare prices and start with $1 in free credit. Model catalog and prices synced 2026-10-06. Generated by `scripts/generate-seo-assets.mjs` from the same copy the site renders — do not edit by hand. Contact: support@bestvirtualgoods.com ## What Y-API is ### Migrate in two lines Keep the OpenAI request and response shapes. openai-python, openai-node, and any SDK, framework, or agent client that speaks the OpenAI protocol can point here directly — your existing code stays untouched. ### Request content never stored The prompts you send and the content the model returns are never saved. Request logs record only time, model, token counts, and credit deducted — for reconciliation, nothing more. ### Top up $1, get $10 in credit The top-up rate is 1:10. Credit is deducted for actual token consumption, and any top-up amount converts at the same rate. You pay for what you call, and nothing when you don't. ### One balance, many keys Keys under the same account share a single balance. Manage them per project or environment, and create or revoke any of them at any time without affecting the rest. ### Every call accounted for The console request log lists the time, model, token counts, credit used, and latency of every call; the overview page aggregates the last 30 days of usage by model. ### No passwords stored Only Google and GitHub sign-in are supported. The site stores no passwords, so there is no attack surface for password leaks or reset flows. ## API configuration - `base_url`: https://api.y-api.bestvirtualgoods.com/v1 - `api_key`: sk-... (create in the console) - `model`: deepseek/deepseek-v4-flash Replace these three items in any OpenAI-compatible SDK. Get the api_key from the "API Keys" page in the console; set model to any ID from "Available models". ## Integration examples ### cURL ```bash curl https://api.y-api.bestvirtualgoods.com/v1/chat/completions \ -H "Authorization: Bearer $YAPI_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek/deepseek-v4-flash", "messages": [{"role": "user", "content": "Hello!"}] }' ``` ### Python (openai SDK) ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.y-api.bestvirtualgoods.com/v1", api_key=os.environ["YAPI_KEY"], # key created in the console ) response = client.chat.completions.create( model="deepseek/deepseek-v4-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ### Node.js (openai SDK) ```javascript import OpenAI from 'openai' const client = new OpenAI({ baseURL: 'https://api.y-api.bestvirtualgoods.com/v1', apiKey: process.env.YAPI_KEY, // key created in the console }) const response = await client.chat.completions.create({ model: 'deepseek/deepseek-v4-flash', messages: [{ role: 'user', content: 'Hello!' }], }) console.log(response.choices[0].message.content) ``` ## Streaming Add the stream parameter — chunks come back in the same SSE format as OpenAI, so clients need no special handling. ```python stream = client.chat.completions.create( model="deepseek/deepseek-v4-flash", messages=[{"role": "user", "content": "Hello!"}], stream=True, ) for chunk in stream: print(chunk.choices[0].delta.content or "", end="") ``` ## Switching models All models share the same base_url and the same key — switching only means changing the model field. The full catalog is on the models page. ```python # cheap, good for high-frequency and batch tasks client.chat.completions.create(model="deepseek/deepseek-v4-flash", messages=msgs) # stronger, good for complex reasoning client.chat.completions.create(model="deepseek/deepseek-v4-flash", messages=msgs) ``` ## Error handling Two errors are the most common, and both leave a record in the console request log. | Status | Meaning | What to do | | --- | --- | --- | | 401 | The key is invalid or has been revoked. | Check the key still exists on the "API Keys" page in the console, or create a new one. | | 403 | Account credit is exhausted. Not 429 — and the message comes back in Chinese, which is why searching for it in English finds nothing. | Top up at https://y-api.bestvirtualgoods.com/app/billing — it recovers immediately, keys stay valid, no code changes needed. | Do not read the cause off the status code: 5 of the 7 codes this gateway returns mean something other than what the number says, and the official SDKs silently retry 2 of those twice before your program ever sees them. The error reference lists every one of them with the message it actually returns. ## Models and prices Every model is called through the same base_url and the same key — switching models only means changing the model field. The model IDs below are ready to copy. | Model ID | Vendor | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | --- | | `deepseek/deepseek-v4-flash` | DeepSeek | $0.00 | $0.00 | | `deepseek/deepseek-v4-pro` | DeepSeek | $0.50 | $1.00 | | `deepseek/deepseek-v4.1-flash` | DeepSeek | $0.20 | $1.00 | | `deepseek/deepseek-v4-flash-0731` | DeepSeek | $0.15 | $0.30 | | `qwen/qwen3.8-flash` | Qwen | $0.20 | $0.50 | | `minimax/minimax-m2.7` | MiniMax | $0.00 | $0.00 | | `z-ai/glm-5.2` | Z.ai | $1.40 | $4.40 | | `z-ai/glm-5.3` | Z.ai | $1.40 | $5.00 | | `z-ai/glm-5.3-flash` | Z.ai | $0.15 | $0.50 | | `moonshotai/kimi-k3` | Moonshot AI | $3.00 | $15.00 | | `moonshotai/kimi-k2.6` | Moonshot AI | $0.95 | $4.00 | | `anthropic/claude-opus-5` | Anthropic | $5.00 | $25.00 | | `anthropic/claude-sonnet-5` | Anthropic | $2.00 | $10.00 | | `openai/gpt-5.6-sol` | OpenAI | $5.00 | $30.00 | | `openai/gpt-5.6-luna` | OpenAI | $0.30 | $1.30 | | `tencent/hy3` | Tencent | $0.00 | $0.00 | | `xiaomi/mimo-v2.5` | Xiaomi | $0.00 | $0.00 | | `tencent/hy4-preview` | Tencent | $1.00 | $3.00 | | `stepfun/step-3.7-flash` | StepFun | $0.20 | $1.20 | | `xiaomi/mimo-v2.6-flash` | Xiaomi | $0.18 | $0.36 | Prices are in USD credit (a $1 top-up adds $10 credit). The catalog and prices may change; the console "Models" page shows the current list at your account's own rate. The table above is quoted in account credit. Divide any figure by 10 for the USD actually paid: `deepseek/deepseek-v4-flash-0731` at $0.15 in credit is $0.015 per 1M input tokens in cash. The cash price is what compares against another provider's list price. Both prices, computed per model, are published as JSON at https://y-api.bestvirtualgoods.com/pricing.json. ## Billing Top up $1 and your account gains $10 in credit; signing up grants $1 in credit with no card required. $1 of sign-up credit ≈ 6.67M input tokens on deepseek/deepseek-v4-flash-0731, the lowest-priced model in the catalog at $0.15 per 1M tokens in credit. Every $1 you top up becomes $10 in credit ≈ 66.7M input tokens on that same model. | You pay | Credit received | | --- | --- | | $5 | $50 | | $10 | $100 | | $50 | $500 | | $100 | $1,000 | Not a plan tier — there is no amount threshold; any top-up converts at the same rate. - **Billing**: Credit is deducted per token of usage. No calls, no charges. - **Minimum spend**: None. An idle account costs nothing either. - **New users**: $1 credit on sign-up, no card required. - **Credit exhausted**: The API returns 403 (insufficient_user_quota). Top up at https://y-api.bestvirtualgoods.com/app/billing — it recovers immediately, and keys stay valid. - **Referral credit**: $1 for each side when an account signs up through another account's invitation link. Sign-up credit still applies, so an invited account starts with $2 in credit and no card. ## Who it fits - **Side projects & solo products**: A demo with no revenue yet shouldn't carry fixed costs. Pay for what you use; pause the project and pay nothing. - **Agents & batch jobs**: Multi-round tool calls, long context, and batch processing multiply token consumption — and the unit-price gap multiplies straight into the bill. - **Internal team tools**: Hand out different keys per environment sharing one balance, and reconcile who called which model, when, from the logs. - **Migrating off the official API**: Your code is already built on the OpenAI SDK, and you don't want to rewrite the call layer just to switch providers. ## Stated up front ### "$1 becomes $10" is credit, not a discount coupon You pay $1 and your account gains $10 in credit. Credit burns at each model's multiplier, which varies widely — so it is not "official pricing at 10% off". Treat the actual deductions in the request log as the real cost. ### Your request content is not stored The prompts you send and the content the model returns pass through in memory and are never written to disk — not by us, and not for training. What we keep is call metadata only: time, model, token counts, latency, and credit deducted, for billing reconciliation. ### Keys can be revoked anytime A key's plaintext is returned on demand only when you click "Reveal"; it never sits exposed in lists. Once deleted, calls using that key fail immediately while other keys are unaffected. A model report you can inspect. The GPT 6 Astra test is public, including the endpoint, results and original source. View model verification: https://y-api.bestvirtualgoods.com/verification ## FAQ ### What is Y-API? Y-API is an OpenAI-compatible LLM API service. It provides a single endpoint and key so the same code can call 20+ models — including DeepSeek, Qwen, MiniMax, Z.ai, Moonshot AI, Anthropic, OpenAI, Tencent, Xiaomi, and StepFun — billed per token. ### Is Y-API compatible with the official OpenAI API? Yes. Y-API uses the same request and response structure as OpenAI. Swap the SDK's base_url for Y-API's endpoint and the api_key for a Y-API key — nothing else changes. ### Which models does Y-API support? Y-API currently supports 20+ models from vendors including DeepSeek, Qwen, MiniMax, Z.ai, Moonshot AI, Anthropic, OpenAI, Tencent, Xiaomi, and StepFun. The full list of model IDs is on the Y-API models page (https://y-api.bestvirtualgoods.com/models); pass any of them in the request's model field. ### Do any models cost nothing at all? Yes — deepseek/deepseek-v4-flash and minimax/minimax-m2.7 and tencent/hy3 and xiaomi/mimo-v2.5 are free through Y-API: no credit is deducted per call, using the same base_url and key as the paid models. The free flag is set by our upstream provider and can be withdrawn; if it is, these calls start billing at the listed rate, with no code change on your side. ### How does Y-API billing work? Account credit is deducted based on token consumption. The deduction depends on the called model's multiplier, with input and output counted separately. The exact deduction for any call is in the console request log. ### What does "top up $1 and get $10 in credit" mean? Every $1 you top up adds $10 to your account credit. Credit is the billing unit inside the service, deducted per token; different models burn it at different rates, so it is not a fixed discount off official pricing. ### Does Y-API have a minimum spend? No. Y-API only deducts credit for actual token usage, and an idle account costs nothing. ### Can I create multiple API keys? Is credit split between them? Yes, you can create multiple keys, and they all share the same account balance — no per-key allocation. This suits per-project or per-environment management; any key can be revoked at any time without affecting the others. ### Does Y-API support streaming? Yes. Add the stream parameter to the request and you get SSE chunks in the same format as OpenAI — no special client handling needed. ### Can I see my call history? Yes. The console request log lists the time, model, key, token counts, credit used, and latency of every call, and the overview page aggregates the last 30 days of usage and request counts by model. ### Can Y-API see the content I send? No. Y-API does not store request or response content. The logs shown in the console contain only call metadata (time, model, token counts, latency, credit deducted) — never prompts or completions. ### What happens when my credit runs out? The API returns HTTP 403 with error code insufficient_user_quota — not 429 — and the request is not executed. Top up at https://y-api.bestvirtualgoods.com/app/billing — credit is available immediately, keys stay valid, and no code changes are needed. ### Do I need a credit card to sign up? No. Sign in with Google or GitHub and your account opens automatically with $1 in credit — no payment details involved during the trial. ### Are there rate limits for trial accounts? After sign-up, you are limited to 500 requests per day (shared across all models, regardless of whether they are free or paid). Once you complete your first top-up (any amount), the limit is removed. Paid models are additionally subject to account balance constraints, with credit deducted based on token usage. ## Pages English is the default locale; mirrors live under /zh (中文), /ja (日本語), /ko (한국어), /es (Español), /pt (Português (BR)), /de (Deutsch). - [DeepSeek — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/deepseek): DeepSeek’s V4 family on Y-API spans four text-first entries: a free V4 Flash preview, the official 0731 Flash release, V4.1 Flash with native image input, and the larger V4 Pro — all with a million-token window and a sparse mixture-of-experts design. ([中文](https://y-api.bestvirtualgoods.com/zh/models/deepseek) [日本語](https://y-api.bestvirtualgoods.com/ja/models/deepseek) [한국어](https://y-api.bestvirtualgoods.com/ko/models/deepseek) [Español](https://y-api.bestvirtualgoods.com/es/models/deepseek) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/deepseek) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/deepseek)) - [Qwen — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/qwen): Alibaba’s Qwen appears on Y-API as a single entry, Qwen3.8 Flash: a multimodal reasoning model that accepts text, images and video over a one-million-token window, with document and chart understanding in the same selection as coding work. ([中文](https://y-api.bestvirtualgoods.com/zh/models/qwen) [日本語](https://y-api.bestvirtualgoods.com/ja/models/qwen) [한국어](https://y-api.bestvirtualgoods.com/ko/models/qwen) [Español](https://y-api.bestvirtualgoods.com/es/models/qwen) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/qwen) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/qwen)) - [MiniMax — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/minimax): MiniMax appears on Y-API with M2.7, a free text-only entry positioned around agentic productivity — debugging, root-cause analysis and multi-step professional work — inside the shortest context window in the catalog at 204,800 tokens. ([中文](https://y-api.bestvirtualgoods.com/zh/models/minimax) [日本語](https://y-api.bestvirtualgoods.com/ja/models/minimax) [한국어](https://y-api.bestvirtualgoods.com/ko/models/minimax) [Español](https://y-api.bestvirtualgoods.com/es/models/minimax) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/minimax) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/minimax)) - [Z.ai — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/z-ai): Z.ai’s GLM 5 family on Y-API has three entries: GLM 5.2 and GLM 5.3 share a million-token text window with different post-training, while GLM 5.3 Flash is a separately trained multimodal base combining sparse and linear attention. ([中文](https://y-api.bestvirtualgoods.com/zh/models/z-ai) [日本語](https://y-api.bestvirtualgoods.com/ja/models/z-ai) [한국어](https://y-api.bestvirtualgoods.com/ko/models/z-ai) [Español](https://y-api.bestvirtualgoods.com/es/models/z-ai) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/z-ai) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/z-ai)) - [Moonshot AI — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/moonshotai): Moonshot AI’s Kimi line on Y-API has two multimodal entries: K2.6 at a 262,144-token window focused on coding and UI work, and K3 at a full million tokens built for long-horizon agents and repository-scale tasks. ([中文](https://y-api.bestvirtualgoods.com/zh/models/moonshotai) [日本語](https://y-api.bestvirtualgoods.com/ja/models/moonshotai) [한국어](https://y-api.bestvirtualgoods.com/ko/models/moonshotai) [Español](https://y-api.bestvirtualgoods.com/es/models/moonshotai) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/moonshotai) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/moonshotai)) - [Anthropic — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/anthropic): Anthropic’s Claude 5 generation on Y-API has two entries — Sonnet 5 and Opus 5 — both accepting text, image and file input over a million-token window. They are the catalog’s reference tier for coding and professional work, and its most expensive alongside GPT-5.6 Sol. ([中文](https://y-api.bestvirtualgoods.com/zh/models/anthropic) [日本語](https://y-api.bestvirtualgoods.com/ja/models/anthropic) [한국어](https://y-api.bestvirtualgoods.com/ko/models/anthropic) [Español](https://y-api.bestvirtualgoods.com/es/models/anthropic) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/anthropic) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/anthropic)) - [OpenAI — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/openai): OpenAI’s GPT-5.6 family on Y-API has two entries with a 1.05M-token window: Sol, the flagship for complex reasoning and command-line coding, and Luna, the efficiency entry for high-volume chat, classification and lightweight agents. ([中文](https://y-api.bestvirtualgoods.com/zh/models/openai) [日本語](https://y-api.bestvirtualgoods.com/ja/models/openai) [한국어](https://y-api.bestvirtualgoods.com/ko/models/openai) [Español](https://y-api.bestvirtualgoods.com/es/models/openai) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/openai) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/openai)) - [Tencent — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/tencent): Tencent’s Hunyuan line on Y-API has two text-only entries: Hy3, a free 295B-parameter MoE with 21B active, and Hy4 preview, which scales to a 770B backbone with 49B active and a million-token window. Both target grounded, tool-driven work. ([中文](https://y-api.bestvirtualgoods.com/zh/models/tencent) [日本語](https://y-api.bestvirtualgoods.com/ja/models/tencent) [한국어](https://y-api.bestvirtualgoods.com/ko/models/tencent) [Español](https://y-api.bestvirtualgoods.com/es/models/tencent) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/tencent) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/tencent)) - [Xiaomi — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/xiaomi): Xiaomi’s MiMo line on Y-API has two low-cost multimodal entries: V2.5, an omnimodal model accepting text, audio, image and video for free, and V2.6 Flash, a reinforcement-learning-tuned efficiency entry at $0.18 input. Both output text only. ([中文](https://y-api.bestvirtualgoods.com/zh/models/xiaomi) [日本語](https://y-api.bestvirtualgoods.com/ja/models/xiaomi) [한국어](https://y-api.bestvirtualgoods.com/ko/models/xiaomi) [Español](https://y-api.bestvirtualgoods.com/es/models/xiaomi) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/xiaomi) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/xiaomi)) - [StepFun — models, pricing & selection | Y-API](https://y-api.bestvirtualgoods.com/models/stepfun): StepFun is represented on Y-API by Step 3.7 Flash, which pairs a sparse language backbone with a vision encoder for image and video input at a 262,144-token window and selectable reasoning effort. It is a single-model lineup here, like Qwen and MiniMax. ([中文](https://y-api.bestvirtualgoods.com/zh/models/stepfun) [日本語](https://y-api.bestvirtualgoods.com/ja/models/stepfun) [한국어](https://y-api.bestvirtualgoods.com/ko/models/stepfun) [Español](https://y-api.bestvirtualgoods.com/es/models/stepfun) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/stepfun) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/stepfun)) - [DeepSeek V4 Flash 0423 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash): DeepSeek V4 Flash 0423 is the earlier, text-only Flash entry in the V4 family. Its sparse mixture-of-experts design targets coding and conversation workloads with a long context window; it is distinct from the later 0731 revision. ([中文](https://y-api.bestvirtualgoods.com/zh/models/deepseek/deepseek-v4-flash) [日本語](https://y-api.bestvirtualgoods.com/ja/models/deepseek/deepseek-v4-flash) [한국어](https://y-api.bestvirtualgoods.com/ko/models/deepseek/deepseek-v4-flash) [Español](https://y-api.bestvirtualgoods.com/es/models/deepseek/deepseek-v4-flash) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/deepseek/deepseek-v4-flash) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/deepseek/deepseek-v4-flash)) - [DeepSeek V4 Pro 0423 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-pro): DeepSeek V4 Pro 0423 is the larger text-reasoning model in the original V4 pair. Its publisher describes a 1.6T-parameter MoE model with 49B activated, designed for demanding reasoning and software work rather than native image understanding. ([中文](https://y-api.bestvirtualgoods.com/zh/models/deepseek/deepseek-v4-pro) [日本語](https://y-api.bestvirtualgoods.com/ja/models/deepseek/deepseek-v4-pro) [한국어](https://y-api.bestvirtualgoods.com/ko/models/deepseek/deepseek-v4-pro) [Español](https://y-api.bestvirtualgoods.com/es/models/deepseek/deepseek-v4-pro) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/deepseek/deepseek-v4-pro) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/deepseek/deepseek-v4-pro)) - [DeepSeek V4.1 Flash — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash): DeepSeek V4.1 Flash adds native image understanding to the Flash line. Its Causal Encoder-Decoder architecture separates input processing from output generation, with a design aimed at input-heavy coding and computer-use workflows. ([中文](https://y-api.bestvirtualgoods.com/zh/models/deepseek/deepseek-v4.1-flash) [日本語](https://y-api.bestvirtualgoods.com/ja/models/deepseek/deepseek-v4.1-flash) [한국어](https://y-api.bestvirtualgoods.com/ko/models/deepseek/deepseek-v4.1-flash) [Español](https://y-api.bestvirtualgoods.com/es/models/deepseek/deepseek-v4.1-flash) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/deepseek/deepseek-v4.1-flash) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/deepseek/deepseek-v4.1-flash)) - [DeepSeek V4 Flash 0731 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731): DeepSeek V4 Flash 0731 is the publisher’s official V4 Flash release, following the earlier preview. It remains a text-only model; its post-training revision focuses on coding, reasoning and agent tasks. ([中文](https://y-api.bestvirtualgoods.com/zh/models/deepseek/deepseek-v4-flash-0731) [日本語](https://y-api.bestvirtualgoods.com/ja/models/deepseek/deepseek-v4-flash-0731) [한국어](https://y-api.bestvirtualgoods.com/ko/models/deepseek/deepseek-v4-flash-0731) [Español](https://y-api.bestvirtualgoods.com/es/models/deepseek/deepseek-v4-flash-0731) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/deepseek/deepseek-v4-flash-0731) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/deepseek/deepseek-v4-flash-0731)) - [Qwen3.8 Flash — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash): Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ([中文](https://y-api.bestvirtualgoods.com/zh/models/qwen/qwen3.8-flash) [日本語](https://y-api.bestvirtualgoods.com/ja/models/qwen/qwen3.8-flash) [한국어](https://y-api.bestvirtualgoods.com/ko/models/qwen/qwen3.8-flash) [Español](https://y-api.bestvirtualgoods.com/es/models/qwen/qwen3.8-flash) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/qwen/qwen3.8-flash) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/qwen/qwen3.8-flash)) - [MiniMax M2.7 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/minimax/minimax-m2.7): MiniMax M2.7 is a text model centered on agentic productivity: debugging, root-cause analysis and multi-step professional work. Its published examples involve external tools; a completion alone does not execute that workflow. ([中文](https://y-api.bestvirtualgoods.com/zh/models/minimax/minimax-m2.7) [日本語](https://y-api.bestvirtualgoods.com/ja/models/minimax/minimax-m2.7) [한국어](https://y-api.bestvirtualgoods.com/ko/models/minimax/minimax-m2.7) [Español](https://y-api.bestvirtualgoods.com/es/models/minimax/minimax-m2.7) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/minimax/minimax-m2.7) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/minimax/minimax-m2.7)) - [GLM 5.2 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.2): GLM 5.2 is Z.ai’s text-reasoning model for long-horizon engineering. Its publisher highlights a million-token context and IndexShare, which reuses sparse-attention indexing work for long inputs. ([中文](https://y-api.bestvirtualgoods.com/zh/models/z-ai/glm-5.2) [日本語](https://y-api.bestvirtualgoods.com/ja/models/z-ai/glm-5.2) [한국어](https://y-api.bestvirtualgoods.com/ko/models/z-ai/glm-5.2) [Español](https://y-api.bestvirtualgoods.com/es/models/z-ai/glm-5.2) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/z-ai/glm-5.2) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/z-ai/glm-5.2)) - [GLM 5.3 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3): GLM 5.3 builds on the GLM 5.2 base model with revised post-training for complex coding and long-running tasks. OpenRouter documents always-on reasoning for this entry, a meaningful distinction when budgeting output. ([中文](https://y-api.bestvirtualgoods.com/zh/models/z-ai/glm-5.3) [日本語](https://y-api.bestvirtualgoods.com/ja/models/z-ai/glm-5.3) [한국어](https://y-api.bestvirtualgoods.com/ko/models/z-ai/glm-5.3) [Español](https://y-api.bestvirtualgoods.com/es/models/z-ai/glm-5.3) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/z-ai/glm-5.3) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/z-ai/glm-5.3)) - [GLM 5.3 Flash — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash): GLM 5.3 Flash introduces native multimodality to the GLM 5 series. Unlike the text-only GLM 5.3 entry, its new base architecture combines sparse and linear attention for visual and long-context workloads. ([中文](https://y-api.bestvirtualgoods.com/zh/models/z-ai/glm-5.3-flash) [日本語](https://y-api.bestvirtualgoods.com/ja/models/z-ai/glm-5.3-flash) [한국어](https://y-api.bestvirtualgoods.com/ko/models/z-ai/glm-5.3-flash) [Español](https://y-api.bestvirtualgoods.com/es/models/z-ai/glm-5.3-flash) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/z-ai/glm-5.3-flash) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/z-ai/glm-5.3-flash)) - [Kimi K3 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k3): Kimi K3 is Moonshot AI’s open-weight multimodal model for coding, knowledge work and long-horizon agents. Its published design emphasizes iterating against repositories, images, logs and runtime feedback rather than generating a single isolated answer. ([中文](https://y-api.bestvirtualgoods.com/zh/models/moonshotai/kimi-k3) [日本語](https://y-api.bestvirtualgoods.com/ja/models/moonshotai/kimi-k3) [한국어](https://y-api.bestvirtualgoods.com/ko/models/moonshotai/kimi-k3) [Español](https://y-api.bestvirtualgoods.com/es/models/moonshotai/kimi-k3) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/moonshotai/kimi-k3) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/moonshotai/kimi-k3)) - [Kimi K2.6 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k2.6): Kimi K2.6 is a native multimodal model focused on coding, UI generation and agent orchestration. The reference catalog lists text and image inputs, with a smaller context window than the K3 entry. ([中文](https://y-api.bestvirtualgoods.com/zh/models/moonshotai/kimi-k2.6) [日本語](https://y-api.bestvirtualgoods.com/ja/models/moonshotai/kimi-k2.6) [한국어](https://y-api.bestvirtualgoods.com/ko/models/moonshotai/kimi-k2.6) [Español](https://y-api.bestvirtualgoods.com/es/models/moonshotai/kimi-k2.6) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/moonshotai/kimi-k2.6) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/moonshotai/kimi-k2.6)) - [Claude Opus 5 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5): Claude Opus 5 is positioned by OpenRouter as Anthropic’s model for demanding reasoning, software engineering and extended agent work. The reference entry lists text, image and file input with text output. ([中文](https://y-api.bestvirtualgoods.com/zh/models/anthropic/claude-opus-5) [日本語](https://y-api.bestvirtualgoods.com/ja/models/anthropic/claude-opus-5) [한국어](https://y-api.bestvirtualgoods.com/ko/models/anthropic/claude-opus-5) [Español](https://y-api.bestvirtualgoods.com/es/models/anthropic/claude-opus-5) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/anthropic/claude-opus-5) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/anthropic/claude-opus-5)) - [Claude Sonnet 5 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/anthropic/claude-sonnet-5): Claude Sonnet 5 is Anthropic’s Sonnet-class model for coding and professional work. The reviewed reference describes adaptive thinking and an updated tokenizer, so migration involves more than changing a display name. ([中文](https://y-api.bestvirtualgoods.com/zh/models/anthropic/claude-sonnet-5) [日本語](https://y-api.bestvirtualgoods.com/ja/models/anthropic/claude-sonnet-5) [한국어](https://y-api.bestvirtualgoods.com/ko/models/anthropic/claude-sonnet-5) [Español](https://y-api.bestvirtualgoods.com/es/models/anthropic/claude-sonnet-5) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/anthropic/claude-sonnet-5) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/anthropic/claude-sonnet-5)) - [GPT-5.6 Sol — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-sol): GPT-5.6 Sol is the flagship entry of the GPT-5.6 family in OpenRouter’s catalog. Its stated focus is complex reasoning, command-line coding and multi-step problem solving, with text, image and file inputs listed in the reference. ([中文](https://y-api.bestvirtualgoods.com/zh/models/openai/gpt-5.6-sol) [日本語](https://y-api.bestvirtualgoods.com/ja/models/openai/gpt-5.6-sol) [한국어](https://y-api.bestvirtualgoods.com/ko/models/openai/gpt-5.6-sol) [Español](https://y-api.bestvirtualgoods.com/es/models/openai/gpt-5.6-sol) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/openai/gpt-5.6-sol) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/openai/gpt-5.6-sol)) - [GPT-5.6 Luna — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-luna): GPT-5.6 Luna is the GPT-5.6 entry aimed at high-volume, latency-sensitive chat, classification and lightweight agent tasks. It is a separate cost-efficiency candidate, not a claim of Sol-equivalent reasoning at a lower price. ([中文](https://y-api.bestvirtualgoods.com/zh/models/openai/gpt-5.6-luna) [日本語](https://y-api.bestvirtualgoods.com/ja/models/openai/gpt-5.6-luna) [한국어](https://y-api.bestvirtualgoods.com/ko/models/openai/gpt-5.6-luna) [Español](https://y-api.bestvirtualgoods.com/es/models/openai/gpt-5.6-luna) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/openai/gpt-5.6-luna) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/openai/gpt-5.6-luna)) - [Hy3 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/tencent/hy3): Tencent Hy3 is a text-only MoE model with 295B total and 21B active parameters in its publisher card. The reference describes direct-answer and reasoning modes, with an emphasis on multi-turn constraints and tool-driven work. ([中文](https://y-api.bestvirtualgoods.com/zh/models/tencent/hy3) [日本語](https://y-api.bestvirtualgoods.com/ja/models/tencent/hy3) [한국어](https://y-api.bestvirtualgoods.com/ko/models/tencent/hy3) [Español](https://y-api.bestvirtualgoods.com/es/models/tencent/hy3) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/tencent/hy3) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/tencent/hy3)) - [MiMo-V2.5 — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.5): Xiaomi MiMo-V2.5 is an omnimodal model whose reference entry lists text, image, audio and video inputs. Its output is text: broad perception support should not be confused with speech or image generation. ([中文](https://y-api.bestvirtualgoods.com/zh/models/xiaomi/mimo-v2.5) [日本語](https://y-api.bestvirtualgoods.com/ja/models/xiaomi/mimo-v2.5) [한국어](https://y-api.bestvirtualgoods.com/ko/models/xiaomi/mimo-v2.5) [Español](https://y-api.bestvirtualgoods.com/es/models/xiaomi/mimo-v2.5) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/xiaomi/mimo-v2.5) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/xiaomi/mimo-v2.5)) - [Hy4 preview — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/tencent/hy4-preview): Tencent Hy4 preview scales the Hy family to a 770B-parameter backbone with 49B active parameters and a longer context than Hy3. It is a text-only preview aimed at coding agents and sustained tool-use workflows. ([中文](https://y-api.bestvirtualgoods.com/zh/models/tencent/hy4-preview) [日本語](https://y-api.bestvirtualgoods.com/ja/models/tencent/hy4-preview) [한국어](https://y-api.bestvirtualgoods.com/ko/models/tencent/hy4-preview) [Español](https://y-api.bestvirtualgoods.com/es/models/tencent/hy4-preview) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/tencent/hy4-preview) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/tencent/hy4-preview)) - [Step 3.7 Flash — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/stepfun/step-3.7-flash): Step 3.7 Flash combines a sparse language backbone with a vision encoder. Its publisher emphasizes perception plus tool orchestration; the OpenRouter entry lists image and video input and selectable reasoning effort. ([中文](https://y-api.bestvirtualgoods.com/zh/models/stepfun/step-3.7-flash) [日本語](https://y-api.bestvirtualgoods.com/ja/models/stepfun/step-3.7-flash) [한국어](https://y-api.bestvirtualgoods.com/ko/models/stepfun/step-3.7-flash) [Español](https://y-api.bestvirtualgoods.com/es/models/stepfun/step-3.7-flash) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/stepfun/step-3.7-flash) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/stepfun/step-3.7-flash)) - [MiMo-V2.6-Flash — API pricing & integration guide | Y-API](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.6-flash): MiMo-V2.6-Flash is Xiaomi’s efficiency-oriented multimodal entry for coding and agent workflows. OpenRouter links the Flash-RL checkpoint card, which describes reinforcement-learning work across varied task environments. ([中文](https://y-api.bestvirtualgoods.com/zh/models/xiaomi/mimo-v2.6-flash) [日本語](https://y-api.bestvirtualgoods.com/ja/models/xiaomi/mimo-v2.6-flash) [한국어](https://y-api.bestvirtualgoods.com/ko/models/xiaomi/mimo-v2.6-flash) [Español](https://y-api.bestvirtualgoods.com/es/models/xiaomi/mimo-v2.6-flash) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models/xiaomi/mimo-v2.6-flash) [Deutsch](https://y-api.bestvirtualgoods.com/de/models/xiaomi/mimo-v2.6-flash)) - [Affordable LLM API, OpenAI-Compatible — Y-API](https://y-api.bestvirtualgoods.com/): OpenAI-compatible LLM API for 20+ models. Pay per token. Compare prices and start with $1 in free credit. ([中文](https://y-api.bestvirtualgoods.com/zh) [日本語](https://y-api.bestvirtualgoods.com/ja) [한국어](https://y-api.bestvirtualgoods.com/ko) [Español](https://y-api.bestvirtualgoods.com/es) [Português (BR)](https://y-api.bestvirtualgoods.com/pt) [Deutsch](https://y-api.bestvirtualgoods.com/de)) - [LLM API models & pricing](https://y-api.bestvirtualgoods.com/models): Compare 20 LLM models, token prices and API integration guides. Find model IDs and choose a model for your OpenAI-compatible application. ([中文](https://y-api.bestvirtualgoods.com/zh/models) [日本語](https://y-api.bestvirtualgoods.com/ja/models) [한국어](https://y-api.bestvirtualgoods.com/ko/models) [Español](https://y-api.bestvirtualgoods.com/es/models) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/models) [Deutsch](https://y-api.bestvirtualgoods.com/de/models)) - [LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing): View LLM API pricing, cash cost per 1M input tokens and billing rules. Pay per token with no minimum spend. ([中文](https://y-api.bestvirtualgoods.com/zh/pricing) [日本語](https://y-api.bestvirtualgoods.com/ja/pricing) [한국어](https://y-api.bestvirtualgoods.com/ko/pricing) [Español](https://y-api.bestvirtualgoods.com/es/pricing) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/pricing) [Deutsch](https://y-api.bestvirtualgoods.com/de/pricing)) - [Daily Free API Key](https://y-api.bestvirtualgoods.com/daily-key): A free API key for everyone: $20 of shared credit every day, reset at 00:05 UTC. Log in, copy the key, call any model in the catalog. ([中文](https://y-api.bestvirtualgoods.com/zh/daily-key) [日本語](https://y-api.bestvirtualgoods.com/ja/daily-key) [한국어](https://y-api.bestvirtualgoods.com/ko/daily-key) [Español](https://y-api.bestvirtualgoods.com/es/daily-key) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/daily-key) [Deutsch](https://y-api.bestvirtualgoods.com/de/daily-key)) - [Integration guide](https://y-api.bestvirtualgoods.com/docs): Point base_url at Y-API and leave the rest of your code unchanged. The examples below run as copied, and all 20 models share the same pattern. ([中文](https://y-api.bestvirtualgoods.com/zh/docs) [日本語](https://y-api.bestvirtualgoods.com/ja/docs) [한국어](https://y-api.bestvirtualgoods.com/ko/docs) [Español](https://y-api.bestvirtualgoods.com/es/docs) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/docs) [Deutsch](https://y-api.bestvirtualgoods.com/de/docs)) - [API error codes](https://y-api.bestvirtualgoods.com/docs/errors): Every failure this gateway returns, measured against the live endpoint on 2026-09-01. Of the 7 failures it returns, 5 do not mean what their status code implies, and 2 of those get retried by the official SDKs before your program ever sees them. ([中文](https://y-api.bestvirtualgoods.com/zh/docs/errors) [日本語](https://y-api.bestvirtualgoods.com/ja/docs/errors) [한국어](https://y-api.bestvirtualgoods.com/ko/docs/errors) [Español](https://y-api.bestvirtualgoods.com/es/docs/errors) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/docs/errors) [Deutsch](https://y-api.bestvirtualgoods.com/de/docs/errors)) - [OpenRouter alternative: Y-API vs OpenRouter](https://y-api.bestvirtualgoods.com/vs/openrouter): Considering an OpenRouter alternative? Compare Y-API's token costs, model coverage, API compatibility and billing terms before you switch. ([中文](https://y-api.bestvirtualgoods.com/zh/vs/openrouter) [日本語](https://y-api.bestvirtualgoods.com/ja/vs/openrouter) [한국어](https://y-api.bestvirtualgoods.com/ko/vs/openrouter) [Español](https://y-api.bestvirtualgoods.com/es/vs/openrouter) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/vs/openrouter) [Deutsch](https://y-api.bestvirtualgoods.com/de/vs/openrouter)) - [Y-API vs the DeepSeek API](https://y-api.bestvirtualgoods.com/vs/deepseek-official): deepseek/deepseek-v4-flash — free through Y-API, no credit deducted, against $0.22–$0.44 per 1M input tokens on DeepSeek's own rate card. Plus where DeepSeek still wins, and what happens if the free flag goes away. ([中文](https://y-api.bestvirtualgoods.com/zh/vs/deepseek-official) [日本語](https://y-api.bestvirtualgoods.com/ja/vs/deepseek-official) [한국어](https://y-api.bestvirtualgoods.com/ko/vs/deepseek-official) [Español](https://y-api.bestvirtualgoods.com/es/vs/deepseek-official) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/vs/deepseek-official) [Deutsch](https://y-api.bestvirtualgoods.com/de/vs/deepseek-official)) - [Y-API vs the official Qwen API](https://y-api.bestvirtualgoods.com/vs/qwen-official): What the Qwen models cost through Y-API versus paying Alibaba Cloud Model Studio directly: cash price per 1M tokens for every Qwen build in the catalog, every Alibaba figure from its cheapest region and shortest input tier. ([中文](https://y-api.bestvirtualgoods.com/zh/vs/qwen-official) [日本語](https://y-api.bestvirtualgoods.com/ja/vs/qwen-official) [한국어](https://y-api.bestvirtualgoods.com/ko/vs/qwen-official) [Español](https://y-api.bestvirtualgoods.com/es/vs/qwen-official) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/vs/qwen-official) [Deutsch](https://y-api.bestvirtualgoods.com/de/vs/qwen-official)) - [Y-API vs Xiaomi MiMo](https://y-api.bestvirtualgoods.com/vs/xiaomi-official): What the same models cost through Y-API versus Xiaomi MiMo: cash price per 1M input tokens for every model both of us sell, each Xiaomi MiMo figure dated 2026-10-02, plus where Xiaomi MiMo is cheaper. ([中文](https://y-api.bestvirtualgoods.com/zh/vs/xiaomi-official) [日本語](https://y-api.bestvirtualgoods.com/ja/vs/xiaomi-official) [한국어](https://y-api.bestvirtualgoods.com/ko/vs/xiaomi-official) [Español](https://y-api.bestvirtualgoods.com/es/vs/xiaomi-official) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/vs/xiaomi-official) [Deutsch](https://y-api.bestvirtualgoods.com/de/vs/xiaomi-official)) - [Y-API vs Together AI](https://y-api.bestvirtualgoods.com/vs/together): What the same models cost through Y-API versus Together AI: cash price per 1M input tokens for every model both of us sell, each Together AI figure dated 2026-08-26, plus where Together AI is cheaper. ([中文](https://y-api.bestvirtualgoods.com/zh/vs/together) [日本語](https://y-api.bestvirtualgoods.com/ja/vs/together) [한국어](https://y-api.bestvirtualgoods.com/ko/vs/together) [Español](https://y-api.bestvirtualgoods.com/es/vs/together) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/vs/together) [Deutsch](https://y-api.bestvirtualgoods.com/de/vs/together)) - [Y-API vs DeepInfra](https://y-api.bestvirtualgoods.com/vs/deepinfra): What the same models cost through Y-API versus DeepInfra: cash price per 1M input tokens for every model both of us sell, each DeepInfra figure dated 2026-08-26, plus where DeepInfra is cheaper. ([中文](https://y-api.bestvirtualgoods.com/zh/vs/deepinfra) [日本語](https://y-api.bestvirtualgoods.com/ja/vs/deepinfra) [한국어](https://y-api.bestvirtualgoods.com/ko/vs/deepinfra) [Español](https://y-api.bestvirtualgoods.com/es/vs/deepinfra) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/vs/deepinfra) [Deutsch](https://y-api.bestvirtualgoods.com/de/vs/deepinfra)) - [LLM API price comparison](https://y-api.bestvirtualgoods.com/compare/prices): The same models priced side by side across vendor official APIs, OpenRouter, Together AI, DeepInfra and Y-API — dollars per 1M tokens, every figure dated and sourced. ([中文](https://y-api.bestvirtualgoods.com/zh/compare/prices) [日本語](https://y-api.bestvirtualgoods.com/ja/compare/prices) [한국어](https://y-api.bestvirtualgoods.com/ko/compare/prices) [Español](https://y-api.bestvirtualgoods.com/es/compare/prices) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/compare/prices) [Deutsch](https://y-api.bestvirtualgoods.com/de/compare/prices)) - [LLM API cost calculator](https://y-api.bestvirtualgoods.com/calculator): Enter a model and your monthly token volume to compare the bill on Y-API, the vendor's own API, OpenRouter, Together AI and DeepInfra — every figure dated. ([中文](https://y-api.bestvirtualgoods.com/zh/calculator) [日本語](https://y-api.bestvirtualgoods.com/ja/calculator) [한국어](https://y-api.bestvirtualgoods.com/ko/calculator) [Español](https://y-api.bestvirtualgoods.com/es/calculator) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/calculator) [Deutsch](https://y-api.bestvirtualgoods.com/de/calculator)) - [Model verification](https://y-api.bestvirtualgoods.com/verification): Inspect the public Hvoy AI report for GPT 6 Astra on Y-API: the tested endpoint, date, seven checks and the scope of the result. ([中文](https://y-api.bestvirtualgoods.com/zh/verification) [日本語](https://y-api.bestvirtualgoods.com/ja/verification) [한국어](https://y-api.bestvirtualgoods.com/ko/verification) [Español](https://y-api.bestvirtualgoods.com/es/verification) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/verification) [Deutsch](https://y-api.bestvirtualgoods.com/de/verification)) - [System status](https://y-api.bestvirtualgoods.com/status): View the current operational status of the Y-API website, public API gateway, and account platform. ([中文](https://y-api.bestvirtualgoods.com/zh/status) [日本語](https://y-api.bestvirtualgoods.com/ja/status) [한국어](https://y-api.bestvirtualgoods.com/ko/status) [Español](https://y-api.bestvirtualgoods.com/es/status) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/status) [Deutsch](https://y-api.bestvirtualgoods.com/de/status)) - [Usage leaderboard](https://y-api.bestvirtualgoods.com/rankings): Public ranking of the busiest Y-API accounts over the last 30 days — sorted by request count, with tokens and credit spent. ([中文](https://y-api.bestvirtualgoods.com/zh/rankings) [日本語](https://y-api.bestvirtualgoods.com/ja/rankings) [한국어](https://y-api.bestvirtualgoods.com/ko/rankings) [Español](https://y-api.bestvirtualgoods.com/es/rankings) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/rankings) [Deutsch](https://y-api.bestvirtualgoods.com/de/rankings)) - [About Y-API](https://y-api.bestvirtualgoods.com/about): An OpenAI-compatible API gateway serving 20 models from DeepSeek, Qwen, MiniMax, Z.ai, Moonshot AI, Anthropic, OpenAI, Tencent, Xiaomi, and StepFun, billed per token. ([中文](https://y-api.bestvirtualgoods.com/zh/about) [日本語](https://y-api.bestvirtualgoods.com/ja/about) [한국어](https://y-api.bestvirtualgoods.com/ko/about) [Español](https://y-api.bestvirtualgoods.com/es/about) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/about) [Deutsch](https://y-api.bestvirtualgoods.com/de/about)) - [Contact & Support](https://y-api.bestvirtualgoods.com/contact): How to reach Y-API support: email response times, what we can help with, and alternative resources for common questions. ([中文](https://y-api.bestvirtualgoods.com/zh/contact) [日本語](https://y-api.bestvirtualgoods.com/ja/contact) [한국어](https://y-api.bestvirtualgoods.com/ko/contact) [Español](https://y-api.bestvirtualgoods.com/es/contact) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/contact) [Deutsch](https://y-api.bestvirtualgoods.com/de/contact)) - [Privacy Policy](https://y-api.bestvirtualgoods.com/privacy): What Y-API collects, what it does not collect, how that information is used and shared, and how you can exercise your rights. ([中文](https://y-api.bestvirtualgoods.com/zh/privacy) [日本語](https://y-api.bestvirtualgoods.com/ja/privacy) [한국어](https://y-api.bestvirtualgoods.com/ko/privacy) [Español](https://y-api.bestvirtualgoods.com/es/privacy) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/privacy) [Deutsch](https://y-api.bestvirtualgoods.com/de/privacy)) - [Terms of Service](https://y-api.bestvirtualgoods.com/terms): Please read before using Y-API: accounts, billing and credit, acceptable use, service availability, and the limits of liability. ([中文](https://y-api.bestvirtualgoods.com/zh/terms) [日本語](https://y-api.bestvirtualgoods.com/ja/terms) [한국어](https://y-api.bestvirtualgoods.com/ko/terms) [Español](https://y-api.bestvirtualgoods.com/es/terms) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/terms) [Deutsch](https://y-api.bestvirtualgoods.com/de/terms)) - [GDPR Notice](https://y-api.bestvirtualgoods.com/gdpr): Supplementary information for users in the EU and EEA: the data controller, the legal bases for processing, your rights, and how to exercise them. ([中文](https://y-api.bestvirtualgoods.com/zh/gdpr) [日本語](https://y-api.bestvirtualgoods.com/ja/gdpr) [한국어](https://y-api.bestvirtualgoods.com/ko/gdpr) [Español](https://y-api.bestvirtualgoods.com/es/gdpr) [Português (BR)](https://y-api.bestvirtualgoods.com/pt/gdpr) [Deutsch](https://y-api.bestvirtualgoods.com/de/gdpr)) ## Detailed model guides ### DeepSeek V4 Flash 0423 DeepSeek V4 Flash 0423 is the earlier, text-only Flash entry in the V4 family. Its sparse mixture-of-experts design targets coding and conversation workloads with a long context window; it is distinct from the later 0731 revision. Model ID: `deepseek/deepseek-v4-flash` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Use it as a baseline for support-ticket classification, text cleanup and small code explanations. Keep a set of difficult examples before deciding whether to move the workload to a newer Flash revision. ##### What to watch for The unsuffixed API ID is not a promise to follow the latest Flash model. OpenRouter labels this entry 0423. Do not assume it gains the image input or behavior of V4.1 Flash. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-04-24 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_completion_tokens`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_a`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate The current catalog marks this model as free: calls do not deduct credit. This does not promise unlimited capacity, permanent availability or a future free price. $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: Free — no credit deducted. Estimated cash equivalent: Free — no credit deducted. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [DeepSeek V4 Flash 0423](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash) | Free — no credit deducted | Free — no credit deducted | | [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) | $0.30 | $0.03 | | [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) | $0.70 | $0.07 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Classify this ticket as billing, login or bug. Give one reason: I was charged twice for the same invoice. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "deepseek/deepseek-v4-flash", "messages": [ { "role": "user", "content": "Classify this ticket as billing, login or bug. Give one reason: I was charged twice for the same invoice." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "deepseek/deepseek-v4-flash", "messages": [ { "role": "user", "content": "Classify this ticket as billing, login or bug. Give one reason: I was charged twice for the same invoice." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Check that the category is billing and the reason uses only the ticket. Add ambiguous and multi-issue tickets; measure wrong routing, not just fluent wording. #### Before you choose ##### Is this the same model as V4 Flash 0731? No. They have separate catalog IDs. The publisher describes 0731 as the official release superseding the preview. Keep the exact ID in saved evaluations and re-test before changing it. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/deepseek/deepseek-v4-flash) - [Publisher model card linked by OpenRouter](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) DeepSeek V4 Flash 0731 is the publisher’s official V4 Flash release, following the earlier preview. It remains a text-only model; its post-training revision focuses on coding, reasoning and agent tasks. ##### [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) DeepSeek V4.1 Flash adds native image understanding to the Flash line. Its Causal Encoder-Decoder architecture separates input processing from output generation, with a design aimed at input-heavy coding and computer-use workflows. ### DeepSeek V4 Pro 0423 DeepSeek V4 Pro 0423 is the larger text-reasoning model in the original V4 pair. Its publisher describes a 1.6T-parameter MoE model with 49B activated, designed for demanding reasoning and software work rather than native image understanding. Model ID: `deepseek/deepseek-v4-pro` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Consider it for architecture reviews that must reconcile several constraints: database transactions, concurrency, failure recovery and migration order. Supply the relevant code and ask for falsifiable failure cases. ##### What to watch for Parameter count is not an end-to-end quality or speed guarantee. This is the 0423 entry, not every later Pro checkpoint; compare its actual fixes with newer Flash and GLM alternatives before allocating a larger budget. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-04-24 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_completion_tokens`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.50 | $1.00 | | Cash equivalent | $0.05 | $0.10 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $1.00. Estimated cash equivalent: $0.10. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [DeepSeek V4 Pro 0423](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-pro) | $1.00 | $0.10 | | [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) | $0.30 | $0.03 | | [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) | $3.90 | $0.39 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Worker A locks account 1 then account 2. Worker B locks account 2 then account 1. Explain the failure and propose a consistent locking rule. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "deepseek/deepseek-v4-pro", "messages": [ { "role": "user", "content": "Worker A locks account 1 then account 2. Worker B locks account 2 then account 1. Explain the failure and propose a consistent locking rule." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "deepseek/deepseek-v4-pro", "messages": [ { "role": "user", "content": "Worker A locks account 1 then account 2. Worker B locks account 2 then account 1. Explain the failure and propose a consistent locking rule." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Require a concrete deadlock interleaving and a globally consistent lock order. Reject answers that merely add retries without addressing the cause. #### Before you choose ##### Does Pro necessarily beat newer Flash revisions? No. The tier name and parameter count do not establish that ranking. Use the same repository task, tool permissions and token budget, then compare test results and total cost. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/deepseek/deepseek-v4-pro) - [Publisher model card linked by OpenRouter](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) DeepSeek V4 Flash 0731 is the publisher’s official V4 Flash release, following the earlier preview. It remains a text-only model; its post-training revision focuses on coding, reasoning and agent tasks. ##### [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) GLM 5.3 builds on the GLM 5.2 base model with revised post-training for complex coding and long-running tasks. OpenRouter documents always-on reasoning for this entry, a meaningful distinction when budgeting output. ### DeepSeek V4.1 Flash DeepSeek V4.1 Flash adds native image understanding to the Flash line. Its Causal Encoder-Decoder architecture separates input processing from output generation, with a design aimed at input-heavy coding and computer-use workflows. Model ID: `deepseek/deepseek-v4.1-flash` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Evaluate it for tasks that connect a screenshot to application code: locating a layout problem, describing a chart before checking its underlying data, or combining visual feedback with a failing test. ##### What to watch for Native vision in a model card does not establish that Y-API forwards your image format correctly. Start with a known image and verify its contents; a plausible response or HTTP 200 is insufficient evidence. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text, Image | | Output | Text | | Listed on OpenRouter | 2026-09-10 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.20 | $1.00 | | Cash equivalent | $0.02 | $0.10 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.70. Estimated cash equivalent: $0.07. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) | $0.70 | $0.07 | | [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) | $0.30 | $0.03 | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try A button is visible at 1440px but clipped at 375px. List three CSS causes and a browser check that can distinguish each one. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "deepseek/deepseek-v4.1-flash", "messages": [ { "role": "user", "content": "A button is visible at 1440px but clipped at 375px. List three CSS causes and a browser check that can distinguish each one." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "deepseek/deepseek-v4.1-flash", "messages": [ { "role": "user", "content": "A button is visible at 1440px but clipped at 375px. List three CSS causes and a browser check that can distinguish each one." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Look for distinguishable checks for fixed widths, overflow and positioning. Then run a separate controlled image test before adding screenshot-based automation. #### Before you choose ##### Can I reuse the V4 Flash integration unchanged? The text request shape is the same, but the ID, pricing and model behavior differ. Adding images introduces another compatibility boundary; test the media representation and task result separately. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/deepseek/deepseek-v4.1-flash) - [Publisher model card linked by OpenRouter](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) DeepSeek V4 Flash 0731 is the publisher’s official V4 Flash release, following the earlier preview. It remains a text-only model; its post-training revision focuses on coding, reasoning and agent tasks. ##### [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ### DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 is the publisher’s official V4 Flash release, following the earlier preview. It remains a text-only model; its post-training revision focuses on coding, reasoning and agent tasks. Model ID: `deepseek/deepseek-v4-flash-0731` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Use the dated ID when you want to compare a specific Flash revision on patch generation or tool-driven maintenance. Preserve the failing test alongside each prompt so regressions are visible. ##### What to watch for A dated model name improves experiment traceability but does not freeze a gateway’s serving configuration. Do not silently replace the earlier unsuffixed ID and treat past results as measurements of this revision. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-07-31 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `parallel_tool_calls`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_a`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.15 | $0.30 | | Cash equivalent | $0.015 | $0.03 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.30. Estimated cash equivalent: $0.03. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) | $0.30 | $0.03 | | [DeepSeek V4 Flash 0423](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash) | Free — no credit deducted | Free — no credit deducted | | [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) | $0.70 | $0.07 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try A Python average function returns sum(xs) / len(xs). Define behavior for an empty list and write two tests before proposing a fix. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "deepseek/deepseek-v4-flash-0731", "messages": [ { "role": "user", "content": "A Python average function returns sum(xs) / len(xs). Define behavior for an empty list and write two tests before proposing a fix." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "deepseek/deepseek-v4-flash-0731", "messages": [ { "role": "user", "content": "A Python average function returns sum(xs) / len(xs). Define behavior for an empty list and write two tests before proposing a fix." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result The response should choose and document an empty-input contract, test it and preserve the nonempty case. Run the tests rather than judging the patch by appearance. #### Before you choose ##### Why keep the 0731 suffix in my request? It selects this catalog entry rather than the earlier Flash entry. Record it with sampling settings and tool versions so later evaluations compare identifiable configurations. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) - [Publisher model card linked by OpenRouter](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [DeepSeek V4 Flash 0423](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash) DeepSeek V4 Flash 0423 is the earlier, text-only Flash entry in the V4 family. Its sparse mixture-of-experts design targets coding and conversation workloads with a long context window; it is distinct from the later 0731 revision. ##### [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) DeepSeek V4.1 Flash adds native image understanding to the Flash line. Its Causal Encoder-Decoder architecture separates input processing from output generation, with a design aimed at input-heavy coding and computer-use workflows. ### Qwen3.8 Flash Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. Model ID: `qwen/qwen3.8-flash` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Try it on a chart-to-explanation workflow: first recover labels, units and values, then calculate a comparison. Keeping extraction separate from interpretation makes visual mistakes easier to detect. ##### What to watch for Video support in the reference catalog does not validate video uploads through Y-API. The linked weights card is named Flash-Next; treat it as architecture background, not proof that the gateway serves that exact checkpoint. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,000,000 tokens | | Input | Text, Image, Video | | Output | Text | | Listed on OpenRouter | 2026-08-26 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logprobs`, `max_tokens`, `presence_penalty`, `reasoning`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.20 | $0.50 | | Cash equivalent | $0.02 | $0.05 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.45. Estimated cash equivalent: $0.045. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | | [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) | $0.40 | $0.04 | | [Step 3.7 Flash](https://y-api.bestvirtualgoods.com/models/stepfun/step-3.7-flash) | $0.80 | $0.08 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Revenue was 120 in Q1 and 150 in Q2, both in USD thousands. Calculate growth and distinguish the percentage change from the absolute increase. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "qwen/qwen3.8-flash", "messages": [ { "role": "user", "content": "Revenue was 120 in Q1 and 150 in Q2, both in USD thousands. Calculate growth and distinguish the percentage change from the absolute increase." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "qwen/qwen3.8-flash", "messages": [ { "role": "user", "content": "Revenue was 120 in Q1 and 150 in Q2, both in USD thousands. Calculate growth and distinguish the percentage change from the absolute increase." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Expect a 25% increase and an absolute change of USD 30,000. For a visual test, supply those numbers in an image without duplicating them in the prompt. #### Before you choose ##### Does a million-token window replace document retrieval? Not automatically. Repeatedly sending whole documents can add cost and dilute evidence. Compare full-document input with retrieved passages on citation accuracy, omissions and total tokens. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/qwen/qwen3.8-flash) - [Publisher model card linked by OpenRouter](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) GLM 5.3 Flash introduces native multimodality to the GLM 5 series. Unlike the text-only GLM 5.3 entry, its new base architecture combines sparse and linear attention for visual and long-context workloads. ##### [Step 3.7 Flash](https://y-api.bestvirtualgoods.com/models/stepfun/step-3.7-flash) Step 3.7 Flash combines a sparse language backbone with a vision encoder. Its publisher emphasizes perception plus tool orchestration; the OpenRouter entry lists image and video input and selectable reasoning effort. ### MiniMax M2.7 MiniMax M2.7 is a text model centered on agentic productivity: debugging, root-cause analysis and multi-step professional work. Its published examples involve external tools; a completion alone does not execute that workflow. Model ID: `minimax/minimax-m2.7` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Use it to turn an incident timeline into testable hypotheses and a diagnostic plan. Provide logs, deployment changes and known constraints instead of asking for a confident diagnosis without evidence. ##### What to watch for Claims about producing documents or spreadsheets describe a tool-enabled system. You still need a runtime for file creation, validation and permission checks; model output is not itself a finished, verified artifact. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 204,800 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-03-18 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `repetition_penalty`, `response_format`, `seed`, `stop`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_p` | #### API pricing & cost estimate The current catalog marks this model as free: calls do not deduct credit. This does not promise unlimited capacity, permanent availability or a future free price. $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: Free — no credit deducted. Estimated cash equivalent: Free — no credit deducted. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [MiniMax M2.7](https://y-api.bestvirtualgoods.com/models/minimax/minimax-m2.7) | Free — no credit deducted | Free — no credit deducted | | [Hy3](https://y-api.bestvirtualgoods.com/models/tencent/hy3) | Free — no credit deducted | Free — no credit deducted | | [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) | $0.30 | $0.03 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Latency rose after a database migration; CPU stayed flat and connection wait time doubled. Give two hypotheses and one test for each. Do not assert a root cause. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "minimax/minimax-m2.7", "messages": [ { "role": "user", "content": "Latency rose after a database migration; CPU stayed flat and connection wait time doubled. Give two hypotheses and one test for each. Do not assert a root cause." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "minimax/minimax-m2.7", "messages": [ { "role": "user", "content": "Latency rose after a database migration; CPU stayed flat and connection wait time doubled. Give two hypotheses and one test for each. Do not assert a root cause." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Separate observations from hypotheses. A useful answer proposes checks for connection contention or changed query behavior and states what evidence would disprove each explanation. #### Before you choose ##### Does M2.7 create Word or Excel files through this endpoint? The text endpoint returns a model response. A tool runner can use that response to create files, but file writing, execution and format validation belong to your application. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/minimax/minimax-m2.7) - [Publisher model card linked by OpenRouter](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Hy3](https://y-api.bestvirtualgoods.com/models/tencent/hy3) Tencent Hy3 is a text-only MoE model with 295B total and 21B active parameters in its publisher card. The reference describes direct-answer and reasoning modes, with an emphasis on multi-turn constraints and tool-driven work. ##### [DeepSeek V4 Flash 0731](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4-flash-0731) DeepSeek V4 Flash 0731 is the publisher’s official V4 Flash release, following the earlier preview. It remains a text-only model; its post-training revision focuses on coding, reasoning and agent tasks. ### GLM 5.2 GLM 5.2 is Z.ai’s text-reasoning model for long-horizon engineering. Its publisher highlights a million-token context and IndexShare, which reuses sparse-attention indexing work for long inputs. Model ID: `z-ai/glm-5.2` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Evaluate it for staged migrations where earlier constraints must survive several steps: schema changes, backfills, dual writes and eventual cleanup. Ask it to identify the rollback boundary at each stage. ##### What to watch for A large context is not persistent memory. Your application must retain the relevant messages and tool results. Keeping every historical log can raise cost without improving the decisions that matter. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-06-16 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `parallel_tool_calls`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $1.40 | $4.40 | | Cash equivalent | $0.14 | $0.44 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $3.60. Estimated cash equivalent: $0.36. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [GLM 5.2](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.2) | $3.60 | $0.36 | | [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) | $3.90 | $0.39 | | [Hy4 preview](https://y-api.bestvirtualgoods.com/models/tencent/hy4-preview) | $2.50 | $0.25 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Plan a zero-downtime rename of a database column used by old and new application versions. Include deployment order, compatibility and rollback points. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "z-ai/glm-5.2", "messages": [ { "role": "user", "content": "Plan a zero-downtime rename of a database column used by old and new application versions. Include deployment order, compatibility and rollback points." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "z-ai/glm-5.2", "messages": [ { "role": "user", "content": "Plan a zero-downtime rename of a database column used by old and new application versions. Include deployment order, compatibility and rollback points." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Look for an expand-and-contract migration, compatibility during mixed versions and a rollback plan before destructive cleanup. Test the order against your actual deployment process. #### Before you choose ##### How does GLM 5.2 differ from GLM 5.3? Z.ai states that 5.3 uses the same base model with changes from post-training. That is a reason to compare behavior on your engineering tasks, not to infer equivalence from the shared base. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/z-ai/glm-5.2) - [Publisher model card linked by OpenRouter](https://huggingface.co/zai-org/GLM-5.2) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) GLM 5.3 builds on the GLM 5.2 base model with revised post-training for complex coding and long-running tasks. OpenRouter documents always-on reasoning for this entry, a meaningful distinction when budgeting output. ##### [Hy4 preview](https://y-api.bestvirtualgoods.com/models/tencent/hy4-preview) Tencent Hy4 preview scales the Hy family to a 770B-parameter backbone with 49B active parameters and a longer context than Hy3. It is a text-only preview aimed at coding agents and sustained tool-use workflows. ### GLM 5.3 GLM 5.3 builds on the GLM 5.2 base model with revised post-training for complex coding and long-running tasks. OpenRouter documents always-on reasoning for this entry, a meaningful distinction when budgeting output. Model ID: `z-ai/glm-5.3` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Consider it for bug fixes that span implementation, tests and a final review. Supply the acceptance conditions first and evaluate whether the patch actually meets them without weakening tests. ##### What to watch for Do not assume a no-thinking switch is available: the reference page says reasoning cannot be disabled. Budget for reasoning and final text, and inspect truncation when the visible answer is unexpectedly empty. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-08-18 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `parallel_tool_calls`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $1.40 | $5.00 | | Cash equivalent | $0.14 | $0.50 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $3.90. Estimated cash equivalent: $0.39. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) | $3.90 | $0.39 | | [GLM 5.2](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.2) | $3.60 | $0.36 | | [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) | $0.40 | $0.04 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try A payment webhook can be delivered twice and out of order. Propose an idempotency design and tests proving it will not credit an account twice. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "z-ai/glm-5.3", "messages": [ { "role": "user", "content": "A payment webhook can be delivered twice and out of order. Propose an idempotency design and tests proving it will not credit an account twice." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "z-ai/glm-5.3", "messages": [ { "role": "user", "content": "A payment webhook can be delivered twice and out of order. Propose an idempotency design and tests proving it will not credit an account twice." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Require a durable uniqueness constraint and an atomic state transition. Include concurrent duplicates and out-of-order delivery in the tests; an in-memory set is not enough. #### Before you choose ##### Can I turn reasoning off on GLM 5.3? The reviewed OpenRouter page says no. It lists selectable effort levels rather than a disabled mode. Y-API parameter forwarding still needs its own compatibility check. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/z-ai/glm-5.3) - [Publisher model card linked by OpenRouter](https://huggingface.co/zai-org/GLM-5.3) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [GLM 5.2](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.2) GLM 5.2 is Z.ai’s text-reasoning model for long-horizon engineering. Its publisher highlights a million-token context and IndexShare, which reuses sparse-attention indexing work for long inputs. ##### [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) GLM 5.3 Flash introduces native multimodality to the GLM 5 series. Unlike the text-only GLM 5.3 entry, its new base architecture combines sparse and linear attention for visual and long-context workloads. ### GLM 5.3 Flash GLM 5.3 Flash introduces native multimodality to the GLM 5 series. Unlike the text-only GLM 5.3 entry, its new base architecture combines sparse and linear attention for visual and long-context workloads. Model ID: `z-ai/glm-5.3-flash` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Try it on interface inspection where screenshots, requirements and code need to agree. Separate factual visual observations from proposed changes so unsupported guesses can be caught. ##### What to watch for Flash is not merely a price setting on GLM 5.3. It is a distinct model with different modalities. Neither native vision nor a tools parameter proves the corresponding Y-API request path works. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text, Image, Video | | Output | Text | | Listed on OpenRouter | 2026-08-26 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `parallel_tool_calls`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.15 | $0.50 | | Cash equivalent | $0.015 | $0.05 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.40. Estimated cash equivalent: $0.04. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) | $0.40 | $0.04 | | [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) | $3.90 | $0.39 | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try A form labels a required field only with a red border. Propose an accessible validation design covering text, focus and screen-reader feedback. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "z-ai/glm-5.3-flash", "messages": [ { "role": "user", "content": "A form labels a required field only with a red border. Propose an accessible validation design covering text, focus and screen-reader feedback." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "z-ai/glm-5.3-flash", "messages": [ { "role": "user", "content": "A form labels a required field only with a red border. Propose an accessible validation design covering text, focus and screen-reader feedback." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Look for a textual error tied to the input, non-color cues and deliberate focus handling. Add a known screenshot to a separate vision test before relying on visual inspection. #### Before you choose ##### Is GLM 5.3 Flash just a smaller GLM 5.3? The publisher describes a newly trained base and a different hybrid attention architecture. Treat it as a separate candidate, not a drop-in quality-equivalent pricing tier. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/z-ai/glm-5.3-flash) - [Publisher model card linked by OpenRouter](https://huggingface.co/zai-org/GLM-5.3-Flash) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) GLM 5.3 builds on the GLM 5.2 base model with revised post-training for complex coding and long-running tasks. OpenRouter documents always-on reasoning for this entry, a meaningful distinction when budgeting output. ##### [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ### Kimi K3 Kimi K3 is Moonshot AI’s open-weight multimodal model for coding, knowledge work and long-horizon agents. Its published design emphasizes iterating against repositories, images, logs and runtime feedback rather than generating a single isolated answer. Model ID: `moonshotai/kimi-k3` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Evaluate it in a repository task with a real feedback loop: locate a failure, propose a patch, run tests and revise. Score completed requirements and unintended changes, not the length of the plan. ##### What to watch for An agent harness supplies tools, state and execution limits; the model does not acquire those through this API call alone. Its larger context also makes repeated full-history requests worth budgeting explicitly. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text, Image, Video | | Output | Text | | Listed on OpenRouter | 2026-07-16 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $3.00 | $15.00 | | Cash equivalent | $0.30 | $1.50 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $10.50. Estimated cash equivalent: $1.05. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Kimi K3](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k3) | $10.50 | $1.05 | | [Kimi K2.6](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k2.6) | $2.95 | $0.295 | | [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) | $17.50 | $1.75 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try An API times out only on large exports. Design a diagnostic sequence that distinguishes slow SQL, serialization cost and proxy timeouts, without changing production data. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "moonshotai/kimi-k3", "messages": [ { "role": "user", "content": "An API times out only on large exports. Design a diagnostic sequence that distinguishes slow SQL, serialization cost and proxy timeouts, without changing production data." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "moonshotai/kimi-k3", "messages": [ { "role": "user", "content": "An API times out only on large exports. Design a diagnostic sequence that distinguishes slow SQL, serialization cost and proxy timeouts, without changing production data." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Require measurements that isolate the three stages and a safe order of investigation. In an agent run, check that each tool result changes the next decision rather than being ignored. #### Before you choose ##### Does calling K3 automatically start multiple agents? No. Multi-agent behavior requires orchestration in your application or client. Set tool permissions, budgets and stopping rules separately from the model selection. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/moonshotai/kimi-k3) - [Publisher model card linked by OpenRouter](https://huggingface.co/moonshotai/Kimi-K3) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Kimi K2.6](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k2.6) Kimi K2.6 is a native multimodal model focused on coding, UI generation and agent orchestration. The reference catalog lists text and image inputs, with a smaller context window than the K3 entry. ##### [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) Claude Opus 5 is positioned by OpenRouter as Anthropic’s model for demanding reasoning, software engineering and extended agent work. The reference entry lists text, image and file input with text output. ### Kimi K2.6 Kimi K2.6 is a native multimodal model focused on coding, UI generation and agent orchestration. The reference catalog lists text and image inputs, with a smaller context window than the K3 entry. Model ID: `moonshotai/kimi-k2.6` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Use it as a candidate for converting interface requirements into components and tests. Include empty, loading, error and narrow-screen states rather than evaluating only an attractive default screenshot. ##### What to watch for The Kimi product’s agent-swarm workflow is not an automatic feature of a Chat Completions response. Image forwarding and tool execution must be validated in the client you actually deploy. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 262,144 tokens | | Input | Text, Image | | Output | Text | | Listed on OpenRouter | 2026-04-20 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `parallel_tool_calls`, `presence_penalty`, `reasoning`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.95 | $4.00 | | Cash equivalent | $0.095 | $0.40 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $2.95. Estimated cash equivalent: $0.295. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Kimi K2.6](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k2.6) | $2.95 | $0.295 | | [Kimi K3](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k3) | $10.50 | $1.05 | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Design the states of a searchable model list: loading, no matches, failed refresh and success. State what happens to existing results when a refresh fails. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "moonshotai/kimi-k2.6", "messages": [ { "role": "user", "content": "Design the states of a searchable model list: loading, no matches, failed refresh and success. State what happens to existing results when a refresh fails." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "moonshotai/kimi-k2.6", "messages": [ { "role": "user", "content": "Design the states of a searchable model list: loading, no matches, failed refresh and success. State what happens to existing results when a refresh fails." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result A failed refresh should not erase usable cached results. Check keyboard access, retry behavior and clear distinction between no results and data not yet loaded. #### Before you choose ##### Should I choose K2.6 or K3 for frontend work? Compare the same component task and visual acceptance checks. K3 has a larger published context; that alone does not determine whether a smaller UI task benefits from it. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/moonshotai/kimi-k2.6) - [Publisher model card linked by OpenRouter](https://huggingface.co/moonshotai/Kimi-K2.6) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Kimi K3](https://y-api.bestvirtualgoods.com/models/moonshotai/kimi-k3) Kimi K3 is Moonshot AI’s open-weight multimodal model for coding, knowledge work and long-horizon agents. Its published design emphasizes iterating against repositories, images, logs and runtime feedback rather than generating a single isolated answer. ##### [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ### Claude Opus 5 Claude Opus 5 is positioned by OpenRouter as Anthropic’s model for demanding reasoning, software engineering and extended agent work. The reference entry lists text, image and file input with text output. Model ID: `anthropic/claude-opus-5` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Consider it as a second-pass reviewer for a consequential architecture change or difficult bug. Give it the original requirements, a proposed diff and concrete failure scenarios to challenge. ##### What to watch for A premium model can still miss a defect or invent an explanation. Validate findings against code and tests. File inputs listed by OpenRouter do not prove that Y-API exposes Anthropic’s native document handling. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,000,000 tokens | | Input | Text, Image, File | | Output | Text | | Listed on OpenRouter | 2026-07-24 | | Parameters listed by OpenRouter | `include_reasoning`, `max_completion_tokens`, `max_tokens`, `reasoning`, `reasoning_effort`, `response_format`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `verbosity` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $5.00 | $25.00 | | Cash equivalent | $0.50 | $2.50 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $17.50. Estimated cash equivalent: $1.75. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) | $17.50 | $1.75 | | [Claude Sonnet 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-sonnet-5) | $7.00 | $0.70 | | [GPT-5.6 Sol](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-sol) | $20.00 | $2.00 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Review this cache policy: private account responses use Cache-Control: public, max-age=300. Explain the risk and propose headers and regression tests. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "anthropic/claude-opus-5", "messages": [ { "role": "user", "content": "Review this cache policy: private account responses use Cache-Control: public, max-age=300. Explain the risk and propose headers and regression tests." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "anthropic/claude-opus-5", "messages": [ { "role": "user", "content": "Review this cache policy: private account responses use Cache-Control: public, max-age=300. Explain the risk and propose headers and regression tests." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result The review should identify cross-user data leakage, propose an appropriate private or no-store policy and test authenticated responses. Avoid accepting generic security advice without a failure mechanism. #### Before you choose ##### Does Opus justify its higher cost than Sonnet? Only a task-level comparison can establish that. Compare defects caught, accepted fixes, retries and total billed tokens. Reserve a more expensive reviewer for cases where it measurably changes the outcome. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/anthropic/claude-opus-5) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Claude Sonnet 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-sonnet-5) Claude Sonnet 5 is Anthropic’s Sonnet-class model for coding and professional work. The reviewed reference describes adaptive thinking and an updated tokenizer, so migration involves more than changing a display name. ##### [GPT-5.6 Sol](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-sol) GPT-5.6 Sol is the flagship entry of the GPT-5.6 family in OpenRouter’s catalog. Its stated focus is complex reasoning, command-line coding and multi-step problem solving, with text, image and file inputs listed in the reference. ### Claude Sonnet 5 Claude Sonnet 5 is Anthropic’s Sonnet-class model for coding and professional work. The reviewed reference describes adaptive thinking and an updated tokenizer, so migration involves more than changing a display name. Model ID: `anthropic/claude-sonnet-5` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Evaluate it for day-to-day pull requests, code explanations and document analysis with explicit acceptance criteria. Keep hard cases for comparison with Opus instead of routing every request to the same tier. ##### What to watch for Do not reuse token-count assumptions from an older Claude tokenizer. Nor should you assume every sampling option is accepted: the reviewed parameter list does not include temperature for this entry. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,000,000 tokens | | Input | Text, Image, File | | Output | Text | | Listed on OpenRouter | 2026-06-30 | | Parameters listed by OpenRouter | `include_reasoning`, `max_completion_tokens`, `max_tokens`, `reasoning`, `reasoning_effort`, `response_format`, `stop`, `structured_outputs`, `tool_choice`, `tools`, `verbosity` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $2.00 | $10.00 | | Cash equivalent | $0.20 | $1.00 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $7.00. Estimated cash equivalent: $0.70. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Claude Sonnet 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-sonnet-5) | $7.00 | $0.70 | | [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) | $17.50 | $1.75 | | [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) | $3.90 | $0.39 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Review a retry loop that retries every HTTP error three times. Explain which error classes should not be retried and propose bounded backoff behavior. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "anthropic/claude-sonnet-5", "messages": [ { "role": "user", "content": "Review a retry loop that retries every HTTP error three times. Explain which error classes should not be retried and propose bounded backoff behavior." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "anthropic/claude-sonnet-5", "messages": [ { "role": "user", "content": "Review a retry loop that retries every HTTP error three times. Explain which error classes should not be retried and propose bounded backoff behavior." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Distinguish transient failures from authentication and validation errors. For writes, require idempotency before retries; check that the proposal limits both attempts and elapsed time. #### Before you choose ##### Can I copy all of my older Claude parameters? Do not assume so. Start with the minimal request, then add each needed option. The tokenizer and the accepted parameter set can differ from earlier models or from a native Anthropic endpoint. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/anthropic/claude-sonnet-5) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) Claude Opus 5 is positioned by OpenRouter as Anthropic’s model for demanding reasoning, software engineering and extended agent work. The reference entry lists text, image and file input with text output. ##### [GLM 5.3](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3) GLM 5.3 builds on the GLM 5.2 base model with revised post-training for complex coding and long-running tasks. OpenRouter documents always-on reasoning for this entry, a meaningful distinction when budgeting output. ### GPT-5.6 Sol GPT-5.6 Sol is the flagship entry of the GPT-5.6 family in OpenRouter’s catalog. Its stated focus is complex reasoning, command-line coding and multi-step problem solving, with text, image and file inputs listed in the reference. Model ID: `openai/gpt-5.6-sol` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Use it as a candidate for repository-level debugging where commands and tests supply feedback. Provide a bounded task, permitted tools and a requirement to report evidence for the final change. ##### What to watch for Choosing the model does not provide a terminal, web search or filesystem. Those belong to the surrounding application. Do not import claims about a hosted agent product into the capabilities of this text request. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,050,000 tokens | | Input | File, Image, Text | | Output | Text | | Listed on OpenRouter | 2026-07-09 | | Parameters listed by OpenRouter | `include_reasoning`, `max_completion_tokens`, `max_tokens`, `reasoning`, `reasoning_effort`, `response_format`, `seed`, `structured_outputs`, `tool_choice`, `tools`, `verbosity` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $5.00 | $30.00 | | Cash equivalent | $0.50 | $3.00 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $20.00. Estimated cash equivalent: $2.00. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [GPT-5.6 Sol](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-sol) | $20.00 | $2.00 | | [GPT-5.6 Luna](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-luna) | $0.95 | $0.095 | | [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) | $17.50 | $1.75 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Tests pass locally but fail in CI with a case-sensitive import error. Give a diagnosis, a minimal fix and a check that prevents recurrence. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "openai/gpt-5.6-sol", "messages": [ { "role": "user", "content": "Tests pass locally but fail in CI with a case-sensitive import error. Give a diagnosis, a minimal fix and a check that prevents recurrence." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "openai/gpt-5.6-sol", "messages": [ { "role": "user", "content": "Tests pass locally but fail in CI with a case-sensitive import error. Give a diagnosis, a minimal fix and a check that prevents recurrence." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Look for a case mismatch in file paths and imports, not a recommendation to disable checks. Verify the fix on a case-sensitive filesystem and include a CI check. #### Before you choose ##### Is this an OpenAI-hosted API endpoint? No. This page documents the Y-API gateway and its exact catalog ID. OpenRouter supplies reference model information; neither that reference nor the name makes the gateway an official OpenAI endpoint. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/openai/gpt-5.6-sol) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [GPT-5.6 Luna](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-luna) GPT-5.6 Luna is the GPT-5.6 entry aimed at high-volume, latency-sensitive chat, classification and lightweight agent tasks. It is a separate cost-efficiency candidate, not a claim of Sol-equivalent reasoning at a lower price. ##### [Claude Opus 5](https://y-api.bestvirtualgoods.com/models/anthropic/claude-opus-5) Claude Opus 5 is positioned by OpenRouter as Anthropic’s model for demanding reasoning, software engineering and extended agent work. The reference entry lists text, image and file input with text output. ### GPT-5.6 Luna GPT-5.6 Luna is the GPT-5.6 entry aimed at high-volume, latency-sensitive chat, classification and lightweight agent tasks. It is a separate cost-efficiency candidate, not a claim of Sol-equivalent reasoning at a lower price. Model ID: `openai/gpt-5.6-luna` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Start with bounded tasks such as intent classification and extracting a small set of fields. Route ambiguous or high-impact cases to review instead of making a cheap model responsible for every decision. ##### What to watch for Model positioning is not a Y-API latency measurement. Structured-output support in the reference also does not remove the need for schema validation, unknown-value handling and an escalation path. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,050,000 tokens | | Input | File, Image, Text | | Output | Text | | Listed on OpenRouter | 2026-07-09 | | Parameters listed by OpenRouter | `include_reasoning`, `max_completion_tokens`, `max_tokens`, `reasoning`, `reasoning_effort`, `response_format`, `seed`, `structured_outputs`, `tool_choice`, `tools`, `verbosity` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.30 | $1.30 | | Cash equivalent | $0.03 | $0.13 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.95. Estimated cash equivalent: $0.095. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [GPT-5.6 Luna](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-luna) | $0.95 | $0.095 | | [GPT-5.6 Sol](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-sol) | $20.00 | $2.00 | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Extract order_id and intent as JSON from: Please cancel order A-1042 before it ships. Use only information present in the sentence. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "openai/gpt-5.6-luna", "messages": [ { "role": "user", "content": "Extract order_id and intent as JSON from: Please cancel order A-1042 before it ships. Use only information present in the sentence." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "openai/gpt-5.6-luna", "messages": [ { "role": "user", "content": "Extract order_id and intent as JSON from: Please cancel order A-1042 before it ships. Use only information present in the sentence." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Validate the two fields and the order ID, then test missing IDs, multiple orders and ambiguous intent. Measure false confident extractions separately from parse failures. #### Before you choose ##### When should a Luna workflow escalate to Sol? Define thresholds from your own evaluations: conflicting evidence, repeated validation failure or a task with substantial reasoning depth. Do not use the model’s self-reported confidence as the sole trigger. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/openai/gpt-5.6-luna) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [GPT-5.6 Sol](https://y-api.bestvirtualgoods.com/models/openai/gpt-5.6-sol) GPT-5.6 Sol is the flagship entry of the GPT-5.6 family in OpenRouter’s catalog. Its stated focus is complex reasoning, command-line coding and multi-step problem solving, with text, image and file inputs listed in the reference. ##### [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ### Hy3 Tencent Hy3 is a text-only MoE model with 295B total and 21B active parameters in its publisher card. The reference describes direct-answer and reasoning modes, with an emphasis on multi-turn constraints and tool-driven work. Model ID: `tencent/hy3` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Try it for document extraction with explicit missing-information handling. Make it distinguish a supplied fact from an assumption, especially when a short business request omits a date or amount. ##### What to watch for A published emphasis on grounded answers is not evidence that hallucination has been eliminated. Include underspecified questions in your evaluation and check how reasoning settings reach the gateway. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 262,144 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-07-06 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `max_completion_tokens`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_p` | #### API pricing & cost estimate The current catalog marks this model as free: calls do not deduct credit. This does not promise unlimited capacity, permanent availability or a future free price. $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: Free — no credit deducted. Estimated cash equivalent: Free — no credit deducted. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Hy3](https://y-api.bestvirtualgoods.com/models/tencent/hy3) | Free — no credit deducted | Free — no credit deducted | | [Hy4 preview](https://y-api.bestvirtualgoods.com/models/tencent/hy4-preview) | $2.50 | $0.25 | | [MiniMax M2.7](https://y-api.bestvirtualgoods.com/models/minimax/minimax-m2.7) | Free — no credit deducted | Free — no credit deducted | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try From this note extract supplier, amount and due_date; use null when absent: Pay Acme Tools USD 240 for the replacement part. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "tencent/hy3", "messages": [ { "role": "user", "content": "From this note extract supplier, amount and due_date; use null when absent: Pay Acme Tools USD 240 for the replacement part." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "tencent/hy3", "messages": [ { "role": "user", "content": "From this note extract supplier, amount and due_date; use null when absent: Pay Acme Tools USD 240 for the replacement part." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Expect the supplied supplier and amount, with a null due date. Add contradictory notes and follow-up corrections to see whether the model preserves the latest constraint. #### Before you choose ##### Does Hy3 always use a reasoning mode? OpenRouter describes a direct no-thinking default plus selectable reasoning modes. Do not assume the Y-API route preserves that default; verify the response and token usage with your intended settings. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/tencent/hy3) - [Publisher model card linked by OpenRouter](https://huggingface.co/tencent/Hy3) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Hy4 preview](https://y-api.bestvirtualgoods.com/models/tencent/hy4-preview) Tencent Hy4 preview scales the Hy family to a 770B-parameter backbone with 49B active parameters and a longer context than Hy3. It is a text-only preview aimed at coding agents and sustained tool-use workflows. ##### [MiniMax M2.7](https://y-api.bestvirtualgoods.com/models/minimax/minimax-m2.7) MiniMax M2.7 is a text model centered on agentic productivity: debugging, root-cause analysis and multi-step professional work. Its published examples involve external tools; a completion alone does not execute that workflow. ### MiMo-V2.5 Xiaomi MiMo-V2.5 is an omnimodal model whose reference entry lists text, image, audio and video inputs. Its output is text: broad perception support should not be confused with speech or image generation. Model ID: `xiaomi/mimo-v2.5` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Evaluate a workflow that reconciles a transcript with written notes, keeping evidence attached to each claim. After the text baseline works, test each media format individually before combining them. ##### What to watch for An omnimodal model behind a gateway may expose only a subset of its input types. Media tokenization and separate charges can also make the text-token estimate incomplete for audio or video workloads. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,050,000 tokens | | Input | Text, Audio, Image, Video | | Output | Text | | Listed on OpenRouter | 2026-04-22 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `max_tokens`, `presence_penalty`, `reasoning`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_p` | #### API pricing & cost estimate The current catalog marks this model as free: calls do not deduct credit. This does not promise unlimited capacity, permanent availability or a future free price. $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: Free — no credit deducted. Estimated cash equivalent: Free — no credit deducted. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [MiMo-V2.5](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.5) | Free — no credit deducted | Free — no credit deducted | | [MiMo-V2.6-Flash](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.6-flash) | $0.36 | $0.036 | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Transcript: We agreed to ship Friday. Notes: Ship Thursday. Summarize the conflict, cite both sources and leave the final date unresolved. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "xiaomi/mimo-v2.5", "messages": [ { "role": "user", "content": "Transcript: We agreed to ship Friday. Notes: Ship Thursday. Summarize the conflict, cite both sources and leave the final date unresolved." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "xiaomi/mimo-v2.5", "messages": [ { "role": "user", "content": "Transcript: We agreed to ship Friday. Notes: Ship Thursday. Summarize the conflict, cite both sources and leave the final date unresolved." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result The answer should preserve the disagreement instead of inventing a resolution. For audio tests, use a recording with a known transcript and check numbers, names and negations. #### Before you choose ##### Can MiMo-V2.5 generate audio from this API? The reviewed catalog lists text output, not audio output. Audio as an input modality does not establish speech synthesis, and Y-API media compatibility requires a separate check. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/xiaomi/mimo-v2.5) - [Publisher model card linked by OpenRouter](https://huggingface.co/XiaomiMiMo/MiMo-V2.5) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [MiMo-V2.6-Flash](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.6-flash) MiMo-V2.6-Flash is Xiaomi’s efficiency-oriented multimodal entry for coding and agent workflows. OpenRouter links the Flash-RL checkpoint card, which describes reinforcement-learning work across varied task environments. ##### [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ### Hy4 preview Tencent Hy4 preview scales the Hy family to a 770B-parameter backbone with 49B active parameters and a longer context than Hy3. It is a text-only preview aimed at coding agents and sustained tool-use workflows. Model ID: `tencent/hy4-preview` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Consider it for a controlled pilot of multi-step engineering tasks. Keep a known baseline, save the complete tool trace and score the final artifact against requirements that were written before the run. ##### What to watch for The preview designation matters for rollout planning. Do not interpret catalog availability as a stable behavior contract; keep a fallback, regression set and a way to halt a run that stops making progress. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,048,576 tokens | | Input | Text | | Output | Text | | Listed on OpenRouter | 2026-08-28 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `max_completion_tokens`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $1.00 | $3.00 | | Cash equivalent | $0.10 | $0.30 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $2.50. Estimated cash equivalent: $0.25. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Hy4 preview](https://y-api.bestvirtualgoods.com/models/tencent/hy4-preview) | $2.50 | $0.25 | | [Hy3](https://y-api.bestvirtualgoods.com/models/tencent/hy3) | Free — no credit deducted | Free — no credit deducted | | [GLM 5.2](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.2) | $3.60 | $0.36 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Plan a resumable import of one million rows. Include checkpoints, duplicate handling, progress reporting and recovery after a worker crash. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "tencent/hy4-preview", "messages": [ { "role": "user", "content": "Plan a resumable import of one million rows. Include checkpoints, duplicate handling, progress reporting and recovery after a worker crash." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "tencent/hy4-preview", "messages": [ { "role": "user", "content": "Plan a resumable import of one million rows. Include checkpoints, duplicate handling, progress reporting and recovery after a worker crash." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Check for durable checkpoints and idempotent writes. Replay a batch after a simulated crash; a convincing plan is not enough if duplicate records appear. #### Before you choose ##### Should Hy4 preview immediately replace Hy3? Not without a controlled comparison. A larger model and context do not guarantee a better result for your workload. Pilot it behind a switch and retain the tested Hy3 path until the evidence supports migration. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/tencent/hy4-preview) - [Publisher model card linked by OpenRouter](https://huggingface.co/tencent/Hy4-preview) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Hy3](https://y-api.bestvirtualgoods.com/models/tencent/hy3) Tencent Hy3 is a text-only MoE model with 295B total and 21B active parameters in its publisher card. The reference describes direct-answer and reasoning modes, with an emphasis on multi-turn constraints and tool-driven work. ##### [GLM 5.2](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.2) GLM 5.2 is Z.ai’s text-reasoning model for long-horizon engineering. Its publisher highlights a million-token context and IndexShare, which reuses sparse-attention indexing work for long inputs. ### Step 3.7 Flash Step 3.7 Flash combines a sparse language backbone with a vision encoder. Its publisher emphasizes perception plus tool orchestration; the OpenRouter entry lists image and video input and selectable reasoning effort. Model ID: `stepfun/step-3.7-flash` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Evaluate it on structured extraction followed by a consistency check: recover chart values or document fields, then verify the arithmetic with an external tool. Separate perception errors from reasoning errors. ##### What to watch for The Flash name does not supply a latency guarantee, and the gateway’s accepted media formats may be narrower than the reference. Verify each reasoning setting rather than assuming the native provider’s configuration maps directly. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 262,144 tokens | | Input | Text, Image, Video | | Output | Text | | Listed on OpenRouter | 2026-05-28 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logprobs`, `max_tokens`, `presence_penalty`, `reasoning`, `reasoning_effort`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.20 | $1.20 | | Cash equivalent | $0.02 | $0.12 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.80. Estimated cash equivalent: $0.08. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [Step 3.7 Flash](https://y-api.bestvirtualgoods.com/models/stepfun/step-3.7-flash) | $0.80 | $0.08 | | [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) | $0.45 | $0.045 | | [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) | $0.40 | $0.04 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try An invoice lists items at 30 and 45, tax of 15 and a total of 100. Check the arithmetic and identify exactly which value is inconsistent. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "stepfun/step-3.7-flash", "messages": [ { "role": "user", "content": "An invoice lists items at 30 and 45, tax of 15 and a total of 100. Check the arithmetic and identify exactly which value is inconsistent." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "stepfun/step-3.7-flash", "messages": [ { "role": "user", "content": "An invoice lists items at 30 and 45, tax of 15 and a total of 100. Check the arithmetic and identify exactly which value is inconsistent." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result The components sum to 90, but the evidence does not show which source field is wrong. A good answer flags the inconsistency without silently changing the invoice. #### Before you choose ##### What should I compare with other Flash models? Use extraction accuracy, schema validity, completed tool calls and total tokens on the same inputs. Measure endpoint latency yourself instead of interpreting Flash as a standardized speed class. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/stepfun/step-3.7-flash) - [Publisher model card linked by OpenRouter](https://huggingface.co/stepfun-ai/Step-3.7-Flash) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [Qwen3.8 Flash](https://y-api.bestvirtualgoods.com/models/qwen/qwen3.8-flash) Qwen3.8 Flash is Alibaba’s multimodal reasoning entry for text, images and video in the OpenRouter catalog. It brings document and chart understanding into the same model selection as coding assistance and long-context analysis. ##### [GLM 5.3 Flash](https://y-api.bestvirtualgoods.com/models/z-ai/glm-5.3-flash) GLM 5.3 Flash introduces native multimodality to the GLM 5 series. Unlike the text-only GLM 5.3 entry, its new base architecture combines sparse and linear attention for visual and long-context workloads. ### MiMo-V2.6-Flash MiMo-V2.6-Flash is Xiaomi’s efficiency-oriented multimodal entry for coding and agent workflows. OpenRouter links the Flash-RL checkpoint card, which describes reinforcement-learning work across varied task environments. Model ID: `xiaomi/mimo-v2.6-flash` Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-06. #### Choosing this model Selection advice by Y-API. The checks below are suggested evaluations, not published test results. ##### Where to start Try it for a longer task that must retain constraints across revisions: a research summary, a code change or a visual review followed by corrections. Track which requirements survive each turn. ##### What to watch for The catalog ID and the linked Flash-RL card use different names. The card is useful technical background, not proof of a gateway’s weights. As with V2.5, listed media inputs do not establish media generation. #### Model specifications These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling. | Field | Value | | --- | --- | | Context window | 1,050,000 tokens | | Input | Text, Image, Video, Audio | | Output | Text | | Listed on OpenRouter | 2026-09-21 | | Parameters listed by OpenRouter | `frequency_penalty`, `include_reasoning`, `logit_bias`, `logprobs`, `max_tokens`, `min_p`, `presence_penalty`, `reasoning`, `repetition_penalty`, `response_format`, `seed`, `stop`, `structured_outputs`, `temperature`, `tool_choice`, `tools`, `top_k`, `top_logprobs`, `top_p` | #### API pricing & cost estimate | Price basis | Input / 1M tokens | Output / 1M tokens | | --- | --- | --- | | Account credit | $0.18 | $0.36 | | Cash equivalent | $0.018 | $0.036 | $1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price. Input tokens / request: 1000. Output tokens / request: 500. Number of requests: 1000. Estimated credit consumed: $0.36. Estimated cash equivalent: $0.036. Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service. ##### Same token budget, different models | Model | Estimated credit consumed | Estimated cash equivalent | | --- | --- | --- | | [MiMo-V2.6-Flash](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.6-flash) | $0.36 | $0.036 | | [MiMo-V2.5](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.5) | Free — no credit deducted | Free — no credit deducted | | [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) | $0.70 | $0.07 | [View LLM API pricing & billing](https://y-api.bestvirtualgoods.com/pricing) #### API integration examples These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted. Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository. ##### A task to try Summarize these constraints and revise the plan without dropping any: no new dependencies; keep the public API; add retry support only for idempotent requests. ```bash curl --fail-with-body --silent --show-error --max-time 120 \ 'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \ -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \ -H 'Content-Type: application/json' \ --data-binary @- <<'JSON' { "model": "xiaomi/mimo-v2.6-flash", "messages": [ { "role": "user", "content": "Summarize these constraints and revise the plan without dropping any: no new dependencies; keep the public API; add retry support only for idempotent requests." } ], "max_tokens": 4096 } JSON ``` ```python import json import os import sys import urllib.error import urllib.request payload = { "model": "xiaomi/mimo-v2.6-flash", "messages": [ { "role": "user", "content": "Summarize these constraints and revise the plan without dropping any: no new dependencies; keep the public API; add retry support only for idempotent requests." } ], "max_tokens": 4096 } request = urllib.request.Request( "https://api.y-api.bestvirtualgoods.com/v1/chat/completions", data=json.dumps(payload).encode("utf-8"), headers={ "Authorization": "Bearer " + os.environ["TOKEN"], "Content-Type": "application/json", }, method="POST", ) try: with urllib.request.urlopen(request, timeout=120) as response: result = json.load(response) print(json.dumps(result, ensure_ascii=False, indent=2)) except urllib.error.HTTPError as error: print(error.read().decode("utf-8"), file=sys.stderr) raise SystemExit(1) ``` ##### How to evaluate the result Check that all three constraints remain explicit after a follow-up change. In a coding run, inspect the dependency diff and test non-idempotent requests to catch silent scope expansion. #### Before you choose ##### Is the Flash-RL model card the exact API identifier? No. Use xiaomi/mimo-v2.6-flash in the request. The linked publisher card names a checkpoint; copying that repository name into the model field would select a different, unlisted identifier. #### Sources & scope Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint. - [OpenRouter model page & specifications](https://openrouter.ai/xiaomi/mimo-v2.6-flash) - [Publisher model card linked by OpenRouter](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) - [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json) #### Models to compare Compare these alternatives on the same inputs. Their descriptions explain different roles; a lower price does not establish equivalent quality. ##### [MiMo-V2.5](https://y-api.bestvirtualgoods.com/models/xiaomi/mimo-v2.5) Xiaomi MiMo-V2.5 is an omnimodal model whose reference entry lists text, image, audio and video inputs. Its output is text: broad perception support should not be confused with speech or image generation. ##### [DeepSeek V4.1 Flash](https://y-api.bestvirtualgoods.com/models/deepseek/deepseek-v4.1-flash) DeepSeek V4.1 Flash adds native image understanding to the Flash line. Its Causal Encoder-Decoder architecture separates input processing from output generation, with a design aimed at input-heavy coding and computer-use workflows.