Claude Haiku 5.5

anthropic/claude-haiku-5.5

Claude Haiku 5.5 is the current small, fast Claude model, aimed at high-volume and cost-sensitive work such as summarization, subagents and browser automation. It carries a 1M-token window at the lowest input price in the Claude lineup here.

Input · Account credit
$0.15

per 1M tokens

Cash equivalent $0.015

Output · Account credit
$0.60

per 1M tokens

Cash equivalent $0.06

Context window
1,000,000

tokens · OpenRouter

Input → Output
Text · Image · File→ Text
Model specifications
Sources reviewed: Y-API catalog snapshot:

Choosing this model

Selection advice by Y-API. The checks below are suggested evaluations, not published test results.

Where to start

Use it as the workhorse behind a larger model: summarizing retrieved context, driving tool calls in a subagent, or handling the bulk of requests that do not need deep reasoning.

What to watch for

A low price per token does not bound the bill when reasoning tokens are billed as output; a long reasoning trace can cost far more than the input suggests. Verify the actual usage in the request log rather than estimating from the headline rate.

Model specifications

These are OpenRouter model-level specifications, not a Y-API compatibility test. A particular route may accept fewer input formats, parameters or tokens. Verify the features you need with a small request; a successful text response does not validate vision or tool calling.

Context window
1,000,000 tokens
Input
Text · Image · File
Output
Text
Listed on OpenRouter
Parameters listed by OpenRouter
include_reasoningmax_completion_tokensmax_tokensreasoningreasoning_effortresponse_formatstopstructured_outputstool_choicetoolsverbosity
OpenRouter model page & specifications

API pricing & cost estimate

Estimated credit consumed
$0.45
Estimated cash equivalent
$0.045

Illustrative token budget, not measured usage or a quote. Includes input and output; excludes separate media charges, cache discounts and retries. Include billed reasoning tokens in your output budget where applicable. Actual billing follows usage returned by the service.

$1 paid adds $10 of credit. Cash equivalent = credit consumed ÷ 10; this is not the model publisher’s list price.

Same token budget, different models
ModelEstimated credit consumedEstimated cash equivalent
Claude Haiku 5.5$0.45$0.045
Claude Haiku 4.5$3.75$0.375
GPT-6 Luna$0.35$0.035
View LLM API pricing & billing

API integration examples

These minimal, text-only requests use the exact Y-API model ID. They do not demonstrate image, audio, video, file or tool support. The output limit also needs room for reasoning; an empty answer with finish_reason=length can mean the budget was exhausted.

A task to try

Summarize this incident report in five bullets for an on-call engineer, keeping every timestamp and affected service name exact.

Set the TOKEN environment variable to a key from your console. Run this on your server or locally; never expose a key in browser code or a public repository.

cURL · Claude Haiku 5.5
curl --fail-with-body --silent --show-error --max-time 120 \
  'https://api.y-api.bestvirtualgoods.com/v1/chat/completions' \
  -H "Authorization: Bearer ${TOKEN:?Set TOKEN first}" \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{
  "model": "anthropic/claude-haiku-5.5",
  "messages": [
    {
      "role": "user",
      "content": "Summarize this incident report in five bullets for an on-call engineer, keeping every timestamp and affected service name exact."
    }
  ],
  "max_tokens": 4096
}
JSON

How to evaluate the result

Check that no timestamp or service name was altered or dropped, then compare against the full-window Haiku 4.5 on the same reports to see which one preserves details at lower cost.

Before you choose

Can Haiku 5.5 replace a larger model for agent work?

For high-volume, well-specified steps, often yes. For multi-step reasoning that has to stay coherent across a long session, benchmark it against the Opus and Sonnet entries on your own tasks before switching traffic.

Sources & scope

Technical facts come from the cited model page and, where available, its linked publisher card. A card describes that checkpoint; it is not proof of the weights a gateway serves. Prices come only from the Y-API catalog. We do not claim measured latency, uptime or benchmark scores for this endpoint.