About this vendor
Vendor-level selection advice by Y-API. The per-model pages linked below carry the model-specific checks; nothing here is a measured comparison.
Where this vendor fits
MiMo is the cheapest way in this catalog to trial audio-input workflows, and V2.6 Flash undercuts even the DeepSeek Flash tier — the entry point when media breadth matters more than frontier reasoning depth.
What to watch for
Omnimodal input upstream does not mean the gateway forwards every media type, and media tokenization can make a text-token cost estimate incomplete. The V2.6 catalog ID and its Flash-RL card name also differ; the card is background, not the served checkpoint.
Models on Y-API
Prices are account-credit snapshots from the Y-API catalog. Each model links to its own guide with a cost estimator and a request example.
| Model | Context window | Input / 1M tokens | Output / 1M tokens |
|---|---|---|---|
MiMo-V2.5xiaomi/mimo-v2.5 | 1,050,000 | Free — no credit deducted | |
MiMo-V2.6-Flashxiaomi/mimo-v2.6-flash | 1,050,000 | $0.18$0.018 | $0.36$0.036 |
Account credit / Cash equivalent · $1 paid adds $10 of credit; the cash equivalent is credit consumed ÷ 10. These are Y-API prices, not the publisher’s list prices.
Choosing within the lineup
Use free V2.5 to establish that a media workflow works at all, then benchmark V2.6 Flash once volume makes efficiency matter. If audio is not involved, compare against Qwen3.8 Flash and GLM 5.3 Flash before assuming MiMo fits.
Before you choose
Do MiMo entries generate audio or images?
No. The catalog lists text output for both entries. Audio and video appear as input modalities only, and media forwarding through Y-API needs its own compatibility test regardless of the model.
Sources & scope
Vendor scope on this site is exactly the reviewed model guides listed above; a vendor page exists only while at least one of its models is reviewed. Claims about architecture or positioning follow the cited pages. We do not publish measured latency, uptime or benchmark scores for any vendor.