Choosing this model
Selection advice by Y-API. The checks below are suggested evaluations, not published test results.
Where to start
Try it for chat, classification and lightweight agent steps where the request volume is high and each individual call is simple enough that a larger model would be wasted on it.
What to watch for
The low input price applies up to a token threshold; above it the listed rate roughly doubles, so very large requests do not scale linearly with the headline number. There is no publisher card to inspect for this closed-weight model.
Model specifications
These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.
- Context window
- 1,050,000 tokens
- Input
- File · Image · Text
- Output
- Text
- Listed on OpenRouter
- Parameters listed by OpenRouter
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatseedstructured_outputstool_choicetoolsverbosity
API pricing & cost estimate
- Estimated credit consumed
- $0.35
- Estimated cash equivalent
- $0.035
Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.
$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.
| Model | Estimated credit consumed | Estimated cash equivalent |
|---|---|---|
| GPT-6 Luna | $0.35 | $0.035 |
| GPT-5.6 Luna | $0.95 | $0.095 |
| Claude Haiku 5.5 | $0.45 | $0.045 |
API integration examples
These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.
A task to try
Route each incoming request to one of three queues based on the text, and explain in one sentence what triggered the choice.
Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.
curl --fail-with-body --silent --show-error --max-time 120 \
'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
-H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
-H 'Content-Type: application/json' \
--data-binary @- <<'JSON'
{
"model": "openai/gpt-6-luna",
"messages": [
{
"role": "user",
"content": "Route each incoming request to one of three queues based on the text, and explain in one sentence what triggered the choice."
}
],
"max_tokens": 4096
}
JSONHow to evaluate the result
Measure routing accuracy on a labeled sample and watch the cost per thousand requests, since the tiered pricing means the average rate depends on your request sizes.
Before you choose
How does Luna relate to GPT-6 Sol?
They are separate catalog entries with separate IDs and prices; Luna is the cheaper, faster tier. Switching between them is a model-field change, but the results are not interchangeable, so re-run your evaluation.
Sources & scope
Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.