# StepFun — models, pricing & selection | Y-API

> StepFun is represented on Y-API by Step 3.7 Flash, which pairs a sparse language backbone with a vision encoder for image and video input at a 262,144-token window and selectable reasoning effort. It is a single-model lineup here, like Qwen and MiniMax.

This is the markdown representation of https://y-api.bestvirtualgoods.com/models/stepfun. Generated by `scripts/generate-seo-assets.mjs` from the same copy the page renders — do not edit by hand.

Vendor scope: 1 reviewed model.

Sources reviewed: 2026-10-04. Y-API catalog snapshot: 2026-10-04.

## About this vendor

Vendor-level selection advice by Y-API. The per-model pages linked below carry the model-specific checks; nothing here is a measured comparison.

### Where this vendor fits

Step 3.7 Flash sits in the Flash price tier with vision, competing directly with Qwen3.8 Flash and GLM 5.3 Flash; its distinct angle is the publisher’s emphasis on perception plus tool orchestration rather than raw context length.

### What to watch for

The Flash name is not a latency class, and the gateway’s accepted media formats may be narrower than the reference entry. With one model there is no in-family fallback — the fallback is another vendor.

## Models on Y-API

Prices are account-credit snapshots from the Y-API catalog. Each model links to its own guide with a cost estimator and a request example.

| Model | Context window | Input / 1M tokens | Output / 1M tokens |
| --- | --- | --- | --- |
| [Step 3.7 Flash](https://y-api.bestvirtualgoods.com/models/stepfun/step-3.7-flash) | 262,144 | $0.20 | $1.20 |

$1 paid adds $10 of credit; the cash equivalent is credit consumed ÷ 10. These are Y-API prices, not the publisher’s list prices.

## Choosing within the lineup

Evaluate it head-to-head with Qwen3.8 Flash and GLM 5.3 Flash on the same extraction tasks, scoring perception errors separately from reasoning errors. If its perceive-then-verify workflow fits, the remaining budget can fund a stronger second-stage model.

## Before you choose

### How does Step 3.7 Flash differ from the other Flash models?

It combines a vision encoder with a sparse language backbone instead of extending a text-first family, and it exposes selectable reasoning effort. Compare extraction accuracy, schema validity and total tokens on identical inputs — the shared Flash name implies none of these.

## Sources & scope

Vendor scope on this site is exactly the reviewed model guides listed above; a vendor page exists only while at least one of its models is reviewed. Claims about architecture or positioning follow the cited pages. We do not publish measured latency, uptime or benchmark scores for any vendor.

- [OpenRouter vendor page](https://openrouter.ai/stepfun)
- [Publisher organization on Hugging Face](https://huggingface.co/stepfun-ai)
- [Y-API catalog & prices (JSON)](https://y-api.bestvirtualgoods.com/models.json)

## Links

- HTML version of this page: https://y-api.bestvirtualgoods.com/models/stepfun
- Site index for agents: https://y-api.bestvirtualgoods.com/llms.txt
- Full reference (single file): https://y-api.bestvirtualgoods.com/llms-full.txt
- OpenAPI 3.1 spec: https://y-api.bestvirtualgoods.com/openapi.json
- Model catalog (JSON, no key needed): https://y-api.bestvirtualgoods.com/models.json
- API base URL: `https://api.y-api.bestvirtualgoods.com/v1`
- Contact: support@bestvirtualgoods.com
- Model catalog: https://y-api.bestvirtualgoods.com/models
- Integration guide: https://y-api.bestvirtualgoods.com/docs
