Qwen3.8 Max

qwen/qwen3.8-max

Qwen3.8 Max is the flagship entry in the Qwen3.8 line: a large mixture-of-experts model with a 1M-token window that accepts text, image and video input. It shares the Qwen3.8 generation with the Flash model but sits a tier above it on price.

Input · Account credit
$2.00

per 1M tokens

Cash equivalent $0.20

Output · Account credit
$6.00

per 1M tokens

Cash equivalent $0.60

Context window
1,000,000

tokens · OpenRouter

Input → Output
Text · Image · Video→ Text
Model specifications
Sources reviewed: Y-API catalog snapshot:

Choosing this model

Selection advice by Y-API. The checks below are suggested evaluations, not published test results.

Where to start

Reserve it for work where the cheaper Flash model has demonstrably failed: long documents that must stay coherent end to end, or multi-step analysis that has to hold several constraints at once.

What to watch for

The listing is a rolling name rather than a dated snapshot, so behavior can shift under the same ID. Ten times the Flash input price is also a real budget decision — verify the gap on your own tasks before moving traffic over.

Model specifications

These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.

Context window
1,000,000 tokens
Input
Text · Image · Video
Output
Text
Listed on OpenRouter
Parameters listed by OpenRouter
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
OpenRouter model page & specifications

API pricing & cost estimate

Estimated credit consumed
$5.00
Estimated cash equivalent
$0.50

Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.

$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.

Same token budget, different models
ModelEstimated credit consumedEstimated cash equivalent
Qwen3.8 Max$5.00$0.50
Qwen3.8 Flash$0.45$0.045
GLM 5.3$3.90$0.39
View LLM API pricing & billing

API integration examples

These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.

A task to try

Read the attached policy document and list every clause that constrains data retention, quoting the sentence that states each one.

Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.

cURL · Qwen3.8 Max
curl --fail-with-body --silent --show-error --max-time 120 \
  'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
  -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{
  "model": "qwen/qwen3.8-max",
  "messages": [
    {
      "role": "user",
      "content": "Read the attached policy document and list every clause that constrains data retention, quoting the sentence that states each one."
    }
  ],
  "max_tokens": 4096
}
JSON

How to evaluate the result

Every quoted clause should appear verbatim in the source and be genuinely about retention. Feed the same document to the Flash model and compare how many clauses each one misses.

Before you choose

Should I start with Max or Flash?

Start with Flash. Move to Max only when you can show a task where Flash fails and Max succeeds on the same inputs; the price difference is large enough that an unverified upgrade is hard to justify.

Sources & scope

Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.