Server tools
Declare tools on routed requests and understand bounded execution, nested model calls, and web search and fetch.
Normally the calling harness owns the tool loop. With server tools, BitRouter advertises selected tools, intercepts their calls, executes them, appends the results, and calls the model again. The caller receives one final response.
Two gates
A server tool runs only when both conditions are true:
- The deployment enables its implementation under
server_tools. - The request declares that tool.
This prevents a configured capability from appearing on every request and prevents a caller from invoking an implementation the operator did not enable.
Enable implementations in Server tools configuration, then opt into them on a request:
{
"tools": [
{"type": "bitrouter:advisor", "args": {"model": "anthropic:claude-opus-4.8"}}
]
}Provider-defined declarations are carried by APIs that preserve provider tools, including Responses and Messages. Chat Completions accepts function tools only; use the bitrouter/fusion model alias when you need the Fusion shorthand on that surface.
Loop bounds
The default loop permits 10 tool rounds, 30 seconds per tool, 120 seconds for the turn, and three consecutive tool-error rounds. Set the deployment-level iteration limit in Server tools configuration.
Reaching a bound stops the loop instead of allowing unbounded model calls. Nested calls, searches, and fetches still consume their normal provider or backend quota.
Tool families
| Family | Implementations | Ownership |
|---|---|---|
| MCP-backed | Tools from configured mcp_servers | The upstream MCP server performs the action |
| Model-backed | Advisor, Sub-agent, Fusion | BitRouter runs nested model calls |
| Web | Web Search, Web Fetch | BitRouter calls configured BYOK backends |
Toolsets keep those implementations provider-agnostic. Each toolset decides whether to advertise on a request and owns calls to its names; tools from different MCP servers are namespaced to avoid collisions.
Enabling a tool is not authorization by itself. Keep virtual-key policy, upstream credentials, network reachability, and the tool's own approval rules scoped to the callers that need them.
Model-backed tools
Model-backed tools make additional model calls inside BitRouter's server-tool loop. Choose by the job the nested call performs:
| Tool | Use it when | Nested result |
|---|---|---|
| Advisor | The parent needs an expert answer to one question | Advice returned to the parent model |
| Sub-agent | The parent can delegate a self-contained task | Worker's final result |
| Fusion | Several models should analyze the same prompt | Judge analysis, then synthesis |
Every nested call has its own tokens, latency, and provider cost. A tool can improve unit-token productivity, but it is not free parallelism.
Enable the implementations
The deployment enables Advisor, Sub-agent, and Fusion in Server tools configuration. Advisor and Sub-agent use the parent model unless the declaration pins another; Fusion inherits deployment defaults unless the declaration overrides them.
Advisor
Declare bitrouter:advisor when the parent should ask a focused question and then continue its own work:
{
"type": "bitrouter:advisor",
"args": {
"model": "anthropic:claude-opus-4.8",
"instructions": "Review only the authentication edge cases."
}
}At call time the parent supplies prompt; it may also override model. The advisor sees the question and configured instructions, not ownership of the outer task.
Sub-agent
Declare bitrouter:subagent for a task whose context and expected result can be stated independently:
{
"type": "bitrouter:subagent",
"args": {
"model": "openai:gpt-5",
"instructions": "Return a concise factual summary."
}
}The parent calls it with task_name and the required task_description. The worker receives the task description, works in isolation, and returns its result. It does not inherit the parent's transcript or environment automatically.
Fusion
Fusion runs a panel of up to eight models in parallel, asks a judge to compare their answers, and lets the parent or a dedicated synthesizer write the final response.
{
"type": "bitrouter:fusion",
"args": {
"panel": [
{"model": "anthropic:claude-opus-4.8"},
{"model": "openai:gpt-5"}
],
"judge": {"model": "anthropic:claude-opus-4.8"}
}
}When server_tools.fusion is configured, bitrouter/fusion is also available as a model alias. The alias injects the deployment defaults and is the simplest cross-protocol entry point.
The judge compares panel answers; it does not silently merge them. Treat Fusion as a deliberate high-cost path for questions where independent perspectives justify the added calls.
Nested tools
Advisor and Sub-agent declarations may include provider-namespaced tools for their nested model. Fusion panel members may do the same. Those declarations do not grant new credentials or bypass provider capability checks.
Use Telemetry to attribute the nested attempts, and Model routing to inspect the models each declaration selects.
Web tools
Web Search and Web Fetch use the same server-tool loop but solve different tasks:
| Tool | Input | Output |
|---|---|---|
bitrouter:web_search | Search query | Normalized result list or backend answer |
bitrouter:web_fetch | URL | Normalized page content |
Both are off by default. The deployment configures ordered backends, and each request explicitly declares the tool.
Configure backends
Select ordered backends and supply their credentials in Server tools configuration. Entries with no resolvable key are skipped; if none resolve, the tool remains disabled. Requests may select an already configured backend, but cannot supply a new credential through a tool declaration.
Web Search also supports a native backend that runs a search-capable model with its provider-native search tool. That path makes a nested model call.
Web Search
{
"type": "bitrouter:web_search",
"args": {"backend": "exa", "max_results": 3}
}The parent model calls web_search with a query. A declaration may pin one configured backend and lower the result cap. A bare or foreign-namespaced web_search remains a provider-native tool; only the explicit bitrouter: namespace selects BitRouter's implementation.
Web Fetch
{
"type": "bitrouter:web_fetch",
"args": {"backend": "firecrawl", "max_content_tokens": 1500}
}The parent calls web_fetch with a URL. max_content_tokens resolves from deployment to declaration and may only be lowered. Exa, Firecrawl, or Tavily fetch the address on their infrastructure; BitRouter sends the URL to the chosen extraction API instead of dereferencing it directly.
Operational boundaries
- Search and fetch content is untrusted model input; keep prompt-injection defenses in the calling workflow.
- Backend requests consume external quota and may leave your network.
- Tool results are bounded by the server-tool loop and appear in request telemetry when export is enabled.
- Backend pinning chooses an already configured implementation; it does not supply or reveal a credential.
How is this guide?