Gemini 3.7 Flash

google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal entry built for fast agentic workflows, coding and multi-step reasoning, with a 1,048,576-token window over text, image, video, file and audio input.

Input · Account credit
$1.50

per 1M tokens

Cash equivalent $0.15

Output · Account credit
$7.50

per 1M tokens

Cash equivalent $0.75

Context window
1,048,576

tokens · OpenRouter

Input → Output
Text · Image · Video · File · Audio→ Text
Model specifications
Sources reviewed: Y-API catalog snapshot:

Choosing this model

Selection advice by Y-API. The checks below are suggested evaluations, not published test results.

Where to start

Use it for tasks that need responsive iteration with reliable multi-step execution: an agent that reads a tool result and immediately decides the next call, or a coding loop that runs many short cycles.

What to watch for

Google deprecated this model on 2026-10-08: requests to it are automatically routed to 3.8 Flash, and its price has been withdrawn from the rate card. The window and the gateway’s media forwarding are unchanged, but what answers a call is no longer this model — pin the model string you actually want. Media input support still needs its own check against the gateway.

Model specifications

These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.

Context window
1,048,576 tokens
Input
Text · Image · Video · File · Audio
Output
Text
Listed on OpenRouter
Parameters listed by OpenRouter
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
OpenRouter model page & specifications

API pricing & cost estimate

Estimated credit consumed
$5.25
Estimated cash equivalent
$0.525

Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.

$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.

Same token budget, different models
ModelEstimated credit consumedEstimated cash equivalent
Gemini 3.7 Flash$5.25$0.525
Gemini 3.8 Flash$5.25$0.525
Gemini 3.5 Flash$8.00$0.80
View LLM API pricing & billing

API integration examples

These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.

A task to try

Given the last three tool outputs, decide the next call and state the stopping condition for this loop.

Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.

cURL · Gemini 3.7 Flash
curl --fail-with-body --silent --show-error --max-time 120 \
  'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
  -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{
  "model": "google/gemini-3.7-flash",
  "messages": [
    {
      "role": "user",
      "content": "Given the last three tool outputs, decide the next call and state the stopping condition for this loop."
    }
  ],
  "max_tokens": 4096
}
JSON

How to evaluate the result

Run the loop on a task with a known solution and count how many steps it takes, whether it terminates on its own, and whether it repeats a failed call instead of changing approach.

Before you choose

Is 3.7 Flash slower or faster than 3.8 Flash?

The question no longer has an answer on this entry: Google deprecated 3.7 Flash on 2026-10-08 and routes its requests to 3.8 Flash, so a call to either name lands on the same model. Call 3.8 Flash directly and measure that.

Sources & scope

Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.