Pricing

Every gateway adds a line to your bill. We take one off.

0% markup on every token, on every model. The savings come from bitrouter/auto choosing the model for each call — not from shaving a percentage off the fee.

usage-based
0%markup · pay-as-you-go

You pay the provider's list price for tokens and nothing to us. Everything the router needs to do its job is included.

Routing across every model on the catalog — never metered
No BitRouter rate limit · no request cap
Unlimited seats · 1 workspace
1M trace receipts / mo · 30-day retention
Evals, guardrails and multi-provider failover
BYOK and self-host — Apache-2.0, no platform fee
Get API key
outcome-based
Customon savings · enterprise

At scale the comparison stops being something you read and becomes something we guarantee — we bill a share of what we save, only on runs that clear your quality bar.

Measured baseline & quality floor
Budget guarantee — never more than we save you
Shared workspaces & spend controls
Retention to 7 years · SSO · SIEM · DPA
Founders + SLA
Talk to the founders
versus other gateways

A markup, a sales call, or a router that lowers the bill.

Picking the provider for a model you chose moves cost by single digits — same model, cheaper host. Picking the modelmoves it by multiples. That is why we don't need a percentage.

BitRouter
OpenRouter
LiteLLM
What it selects for you
Model × provider, per call
Provider, for a model you pick
Neither — you configure routes
Which cost lever that moves
Model choice — multiples
Host choice — single digits
None — your config decides
Gateway fee on tokens
0% markup
5.5% to load credits
None on the OSS proxy
Hosted tier
Self-serve, no call
Self-serve
Quote-only, sales call
BYOK fee
None
5% past the free allowance
n/a — self-hosted
Self-host
Apache-2.0, free
n/a
MIT, free
SSO · audit logs · RBAC
Enterprise
n/a
Enterprise

Competitor figures are their published rates at time of writing; LiteLLM Enterprise is quote-only, so there is no rate to compare. See the OpenRouter and LiteLLM comparisons for the detail.

your numbers, not ours

You set the target. We report against it.

We'd rather show you your own numbers than a projection of them. Each workload declares what it is optimizing for, and every session is measured against that. Routing you can't hold to a number is just a black box with opinions.

Axis
What you define
What we report
Cost
A budget or cost-per-run ceiling for the workload
Actual spend per session against that ceiling, straight from the receipts — plus what the same sessions would have cost on your baseline model.
Latency
The p50 / p95 you need the route to hold
End-to-end latency per request measured against it, so you can see what a cheaper route cost you in time before you keep it.
Accuracy
The quality floor a route has to clear
Pass rate against that floor. Success rate is the default and needs nothing from you; point an eval at it when your bar is domain-specific.

None of the three costs extra, and none of them waits on us. Success rate ships as the default quality metric — outcome classification is deterministic, with no judge in the request path — so a route has to earn its traffic before it keeps it. An eval only refines that bar where your definition of good is narrower than ours. Beyond it sits the enterprise engagement, where we measure the baseline with you and price on the savings. For measured runs against an all-frontier baseline, see the routed benchmark.

faq

Questions.

OpenRouter picks the best provider for a model you chose, and charges 5.5% to load credits. LiteLLM doesn't pick anything — you configure the routes yourself, and its hosted tier is quote-only through a sales call. BitRouter picks the model and the provider for each call, based on where it sits in your agent's trajectory and how much risk it carries, and charges 0%. That difference isn't a discount, it's a different lever: choosing a cheaper host for the same model moves cost by single digits, while choosing a cheaper model for the calls that don't need a frontier one moves it by multiples. We don't need a percentage because the routing is the product.
Tokens, at the provider's published list price, with 0% markup — no routing fee, no platform fee, no seat fee. Routing itself is never metered, and that's a commitment rather than a not-yet: your request volume will not turn into a billed line later. A free account also includes 1M trace receipts a month kept 30 days, evals, guardrails, multi-provider failover, unlimited seats and one workspace. Point BitRouter at your own provider contracts and we take no percentage on that traffic either — unlike gateways that levy one on BYOK — or self-host the whole Apache-2.0 stack, with the same routing engine, guardrails and observability as the hosted edge, and owe us nothing at all.
Neither — it's a unit of account. Cost-per-session is how you compare bitrouter/auto against running a frontier model outright, in the unit the router actually controls; your invoice is tokens at list price. And we deliberately don't quote it in advance, because an accurate forecast would have to absorb the variance in your context shape, how often the router escalates, and upstream prices that move — the spread we'd charge to cover that would be exactly the markup we removed. What we can do is show you what your traffic did cost, computed from your own receipts, against what your baseline model would have cost for the same sessions. At enterprise scale the variance does get absorbed: that's what the budget guarantee is.
Because a route has to earn its traffic. You declare what each workload is optimizing for — a cost ceiling, a p50/p95 target, a quality floor — and every session is measured against it. Success rate is the default quality metric and needs nothing from you: outcome classification is deterministic, with no judge in the request path, so a failure escalates a route immediately and a cheaper route must succeed repeatedly before it earns traffic. If your definition of good is narrower than that, point an eval at it — `bitrouter optimize` then runs your workflow twice, once as-is and once with a single routing change, and reports the cost and quality deltas so you can publish or roll back.
A free account keeps receipts for 30 days, sized for debugging rather than for an audit. Longer windows are set by the regimes that actually govern log retention rather than by round numbers: 6 months matches the EU AI Act's Article 19 minimum for providers of high-risk AI systems, applicable since 2 August 2026; 12 months matches SOC 2 expectations and PCI DSS 4.0; 7 years covers HIPAA and SOX. Whether a given obligation applies to your system is a call for your own counsel — under the AI Act the duty sits with the provider of the high-risk system, not with BitRouter. Either way we store receipts, never prompts or responses.
You run your full production loop through BitRouter. You set a budget and a measurable quality floor. We guarantee the loop stays under your budget, and we bill a custom share of what we save you against your measured baseline — only on runs that clear your quality bar, and never more than we saved you. It's enterprise-only, because agreeing a baseline and a quality bar takes a conversation; talk to the founders to scope the rate.