Choosing this model
Selection advice by Y-API. The checks below are suggested evaluations, not published test results.
Where to start
Try it on a chart-to-explanation workflow: first recover labels, units and values, then calculate a comparison. Keeping extraction separate from interpretation makes visual mistakes easier to detect.
What to watch for
Video support in the reference catalog does not validate video uploads through Y-API. The linked weights card is named Flash-Next; treat it as architecture background, not proof that the gateway serves that exact checkpoint.
Model specifications
These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.
- Context window
- 1,000,000 tokens
- Input
- Text · Image · Video
- Output
- Text
- Listed on OpenRouter
- Parameters listed by OpenRouter
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Estimate your cost
- Estimated credit consumed
- $0.45
- Estimated cash equivalent
- $0.045
Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.
$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.
| Model | Estimated credit consumed | Estimated cash equivalent |
|---|---|---|
| Qwen3.8 Flash | $0.45 | $0.045 |
| GLM 5.3 Flash | $0.40 | $0.04 |
| Step 3.7 Flash | $0.80 | $0.08 |
Use the API
These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.
A task to try
Revenue was 120 in Q1 and 150 in Q2, both in USD thousands. Calculate growth and distinguish the percentage change from the absolute increase.
Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.
curl --fail-with-body --silent --show-error --max-time 120 \
'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
-H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
-H 'Content-Type: application/json' \
--data-binary @- <<'JSON'
{
"model": "qwen/qwen3.8-flash",
"messages": [
{
"role": "user",
"content": "Revenue was 120 in Q1 and 150 in Q2, both in USD thousands. Calculate growth and distinguish the percentage change from the absolute increase."
}
],
"max_tokens": 4096
}
JSONHow to evaluate the result
Expect a 25% increase and an absolute change of USD 30,000. For a visual test, supply those numbers in an image without duplicating them in the prompt.
Before you choose
Does a million-token window replace document retrieval?
Not automatically. Repeatedly sending whole documents can add cost and dilute evidence. Compare full-document input with retrieved passages on citation accuracy, omissions and total tokens.
Sources & scope
Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.