GPT-5.6 Luna

openai/gpt-5.6-luna

GPT-5.6 Luna is the GPT-5.6 entry aimed at high-volume, latency-sensitive chat, classification and lightweight agent tasks. It is a separate cost-efficiency candidate, not a claim of Sol-equivalent reasoning at a lower price.

Input · Account credit
$0.30

per 1M tokens

Cash equivalent $0.03

Output · Account credit
$1.30

per 1M tokens

Cash equivalent $0.13

Context window
1,050,000

tokens · OpenRouter

Input → Output
File · Image · Text→ Text
Model specifications
Sources reviewed: Y-API catalog snapshot:

Choosing this model

Selection advice by Y-API. The checks below are suggested evaluations, not published test results.

Where to start

Start with bounded tasks such as intent classification and extracting a small set of fields. Route ambiguous or high-impact cases to review instead of making a cheap model responsible for every decision.

What to watch for

Model positioning is not a Y-API latency measurement. Structured-output support in the reference also does not remove the need for schema validation, unknown-value handling and an escalation path.

Model specifications

These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.

Context window
1,050,000 tokens
Input
File · Image · Text
Output
Text
Listed on OpenRouter
Parameters listed by OpenRouter
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetoolsverbosity
OpenRouter model page & specifications

Estimate your cost

Estimated credit consumed
$0.95
Estimated cash equivalent
$0.095

Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.

$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.

Same token budget, different models
ModelEstimated credit consumedEstimated cash equivalent
GPT-5.6 Luna$0.95$0.095
GPT-5.6 Sol$20.00$2.00
Qwen3.8 Flash$0.45$0.045

Use the API

These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.

A task to try

Extract order_id and intent as JSON from: Please cancel order A-1042 before it ships. Use only information present in the sentence.

Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.

cURL · GPT-5.6 Luna
curl --fail-with-body --silent --show-error --max-time 120 \
  'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
  -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{
  "model": "openai/gpt-5.6-luna",
  "messages": [
    {
      "role": "user",
      "content": "Extract order_id and intent as JSON from: Please cancel order A-1042 before it ships. Use only information present in the sentence."
    }
  ],
  "max_tokens": 4096
}
JSON

How to evaluate the result

Validate the two fields and the order ID, then test missing IDs, multiple orders and ambiguous intent. Measure false confident extractions separately from parse failures.

Before you choose

When should a Luna workflow escalate to Sol?

Define thresholds from your own evaluations: conflicting evidence, repeated validation failure or a task with substantial reasoning depth. Do not use the model’s self-reported confidence as the sole trigger.

Sources & scope

Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.