Supported Models

The models hosted BitRouter routes to natively — a curated SOTA set for agentic and coding work, with input, cached-input, and output pricing — plus how to bring your own.

6 min readEdit this page

These are the models hosted BitRouter supports natively for context-aware routing: the ones the router can score, downgrade, and fail over between without any setup on your side. It is a curated set, not an everything-catalog — we track the best SOTA models for agentic and coding workloads, so the router always has a strong frontier option and a strong cheap option for every step of a loop.

You are never limited to this list. Any model you run yourself — local or remote, a fine-tune, a private cluster — routes alongside these with the same ids, fallbacks, and variants: see bring your own model. To reach a model on your own provider account instead of ours, see bring your own provider.

Model catalog

ModelContextModalitiesOpen weightsInput $/MCached input $/MOutput $/M
anthropic/claude-fable-51Mtext, image$10$1$50
anthropic/claude-haiku-4.5200Ktext, image$1$0.1$5
anthropic/claude-opus-4.6200Ktext, image$5$0.5$25
anthropic/claude-opus-4.7200Ktext, image$4.5$0.45$22.5
anthropic/claude-opus-4.81Mtext, image$5$0.5$25
anthropic/claude-opus-51Mtext, image$5$0.5$25
anthropic/claude-sonnet-4.61Mtext, image$3$0.3$15
anthropic/claude-sonnet-51Mtext, image$2$0.2$10
deepseek/deepseek-v3.2128Ktext$0.21$0.105$0.315
deepseek/deepseek-v4-flash256Ktext$0.105$0.0021$0.21
deepseek/deepseek-v4-flash-07311Mtext$0.14$0.028$0.28
deepseek/deepseek-v4-pro256Ktext$1.305$0.10875$2.61
deepseek/deepseek-v4-pro-08131Mtext$0.435$0.003625$0.87
google/gemini-3.1-flash-lite-preview1Mtext, image$0.25$0.025$1.5
google/gemini-3.1-pro-preview2Mtext, image$2$0.2$12
google/gemini-3.5-flash1Mtext, image, audio$1.5$0.15$9
google/gemini-3.6-flash1Mtext, image, audio, video$1.5$0.15$7.5
minimax/minimax-m2.7192Ktext$0.225$0.045$0.9
minimax/minimax-m31Mtext, image$0.3$0.06$1.2
moonshotai/kimi-k2.6256Ktext$0.7125$0.12$3
moonshotai/kimi-k2.7-code256Ktext, image$0.7125$0.1425$3
moonshotai/kimi-k31Mtext, image$2$6
openai/gpt-5.4128Ktext, image$2.5$0.25$15
openai/gpt-5.4-mini128Ktext, image$0.75$0.075$4.5
openai/gpt-5.5128Ktext, image$5$0.5$30
openai/gpt-5.6-luna400Ktext, image$0.2$0.02$1.2
openai/gpt-5.6-sol1Mtext, image$5$0.5$30
openai/gpt-5.6-terra1Mtext, image$2$0.2$12
qwen/qwen3.5-122b-a10b256Ktext, image$0.26$2.08
qwen/qwen3.5-27b256Ktext, image$0.25$2
qwen/qwen3.6-27b256Ktext, image$0.3$3.2
qwen/qwen3.6-35b-a3b256Ktext, image$0.248$1.485
qwen/qwen3.6-flash1Mtext, image$0.1875$0.01875$1.125
qwen/qwen3.7-max1Mtext$1.875$0.375$5.625
qwen/qwen3.7-plus1Mtext, image$0.4$0.08$1.6
qwen/qwen3.8-2.4t-a95b1Mtext$2$0.17$6
qwen/qwen3.8-27b1Mtext, image, video$0.5$0.05$3
qwen/qwen3.8-max1Mtext, image, video$2$0.25$6
x-ai/grok-4.20128Ktext, image$1.25$0.2$2.5
x-ai/grok-4.20-multi-agent1Mtext, image$1.25$0.2$2.5
x-ai/grok-4.31Mtext, image$1.25$0.2$2.5
x-ai/grok-4.5500Ktext, image$2$0.3$6
x-ai/grok-4.6500Ktext, image$2$0.5$6
x-ai/grok-build-0.1256Ktext$1$0.2$2
xiaomi/mimo-v2.5256Ktext$0.14$0.0028$0.28
xiaomi/mimo-v2.5-pro256Ktext$0.435$0.0036$0.87
z-ai/glm-4.7200Ktext$0.6$0.11$2.2
z-ai/glm-5.1128Ktext$0.98$0.182$3.08
z-ai/glm-5.21Mtext$0.979$0.182$3.08
z-ai/glm-5.31Mtext$1.4$0.26$4.4
z-ai/glm-5.3-flash1Mtext, image, video$0.075$0.015$0.25

Prices are USD per million tokens — what a hosted BitRouter request costs today, from the cheapest provider actually serving that model ( means no per-token provider is currently serving it; cached input is where the model prices no cache tier). Modalities are the inputs a model accepts; every model returns text. A provider listed in the public, open-source registry but not currently reachable isn't priced here, so bringing your own key to one of them can beat these rates — and anyone can register a provider by opening a PR against it.

How model ids work

On BitRouter a "model" is not a single endpoint. It's an aggregate: one logical model — say openai/gpt-4o or anthropic/claude-sonnet-4.6 — that can be served by many providers at once. You address it by the stable model id in the catalog above, and BitRouter decides which underlying provider endpoint actually answers each request. That indirection is the point: you write your agent against anthropic/claude-sonnet-4.6, and the set of providers behind it can grow, shrink, or re-price without you changing a line of code.

Because a model is an aggregate, requesting one kicks off a provider selection step — BitRouter ranks the eligible providers on a blend of cost, latency, throughput, and uptime, and sends your request to the best one. A transient failure falls through to the next-ranked provider, or to the next model you listed via fallback. Append a variant suffix — :cost, :latency, :throughput — to re-rank the eligible providers along one axis for a single request; a bare id is the balanced default.

Calling these models

Every model above is reachable on hosted BitRouter — one account against api.bitrouter.ai, no upstream provider keys, no per-provider signups. You pay BitRouter directly at the prices listed here, billed per request; failed requests aren't billed.

bitrouter cloud login   # one-time device-flow sign-in
bitrouter start         # the `bitrouter` provider auto-enables once signed in

How is this guide?

On this page