Choosing this model
Selection advice by Y-API. The checks below are suggested evaluations, not published test results.
Where to start
Try it as the default Flash-tier workhorse for coding assistance and agentic steps, where you want the current generation without paying the Pro tier’s per-token rate.
What to watch for
This is the entry the older Flash generations now route to: Google deprecated 3.7 Flash on 2026-10-08 and sends its requests here, so a migration is about picking the right model string rather than about budget. The very wide modality list is a model specification; each format still needs a gateway compatibility check.
Model specifications
These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.
- Context window
- 1,048,576 tokens
- Input
- Text · Image · Video · File · Audio
- Output
- Text
- Listed on OpenRouter
- Parameters listed by OpenRouter
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
API pricing & cost estimate
- Estimated credit consumed
- $5.25
- Estimated cash equivalent
- $0.525
Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.
$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.
| Model | Estimated credit consumed | Estimated cash equivalent |
|---|---|---|
| Gemini 3.8 Flash | $5.25 | $0.525 |
| Gemini 3.7 Flash | $5.25 | $0.525 |
| Qwen3.8 Flash | $0.45 | $0.045 |
API integration examples
These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.
A task to try
Review this diff for correctness and say which change you would revert first if the test suite fails after merging.
Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.
curl --fail-with-body --silent --show-error --max-time 120 \
'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
-H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
-H 'Content-Type: application/json' \
--data-binary @- <<'JSON'
{
"model": "google/gemini-3.8-flash",
"messages": [
{
"role": "user",
"content": "Review this diff for correctness and say which change you would revert first if the test suite fails after merging."
}
],
"max_tokens": 4096
}
JSONHow to evaluate the result
Check whether the identified change is genuinely the riskiest one by reverting it in a scratch branch and running the tests. Then compare the review against the 3.7 entry on the same diff.
Before you choose
Which Gemini Flash should new integrations start with?
Start with the newest, 3.8 Flash — it is also what calls to the deprecated 3.7 Flash now reach, so a request to either name lands on the same model. Keep the evaluation harness so you can re-check when the next Flash entry arrives.
Sources & scope
Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.