Choosing this model
Selection advice by Y-API. The checks below are suggested evaluations, not published test results.
Where to start
Consider it for bug fixes that span implementation, tests and a final review. Supply the acceptance conditions first and evaluate whether the patch actually meets them without weakening tests.
What to watch for
Do not assume a no-thinking switch is available: the reference page says reasoning cannot be disabled. Budget for reasoning and final text, and inspect truncation when the visible answer is unexpectedly empty.
Model specifications
These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.
- Context window
- 1,048,576 tokens
- Input
- Text
- Output
- Text
- Listed on OpenRouter
- Parameters listed by OpenRouter
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Estimate your cost
- Estimated credit consumed
- $3.90
- Estimated cash equivalent
- $0.39
Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.
$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.
| Model | Estimated credit consumed | Estimated cash equivalent |
|---|---|---|
| GLM 5.3 | $3.90 | $0.39 |
| GLM 5.2 | $3.60 | $0.36 |
| GLM 5.3 Flash | $0.40 | $0.04 |
Use the API
These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.
A task to try
A payment webhook can be delivered twice and out of order. Propose an idempotency design and tests proving it will not credit an account twice.
Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.
curl --fail-with-body --silent --show-error --max-time 120 \
'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
-H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
-H 'Content-Type: application/json' \
--data-binary @- <<'JSON'
{
"model": "z-ai/glm-5.3",
"messages": [
{
"role": "user",
"content": "A payment webhook can be delivered twice and out of order. Propose an idempotency design and tests proving it will not credit an account twice."
}
],
"max_tokens": 4096
}
JSONHow to evaluate the result
Require a durable uniqueness constraint and an atomic state transition. Include concurrent duplicates and out-of-order delivery in the tests; an in-memory set is not enough.
Before you choose
Can I turn reasoning off on GLM 5.3?
The reviewed OpenRouter page says no. It lists selectable effort levels rather than a disabled mode. Y-API parameter forwarding still needs its own compatibility check.
Sources & scope
Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.