Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
crossbar
crossbar is an affinity router in front of several llama-server routers. A client's identity is
the first path segment of its base URL — its route. Each conversation takes a sticky lease on one
host, chosen for the most free slots for its model times weight, and streams the answer back
incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all
new in v1.
Build
Build everything with make build; the binaries land in bin/. Check the work with make gate,
which runs the formatter, vet, tests and line-length check with no network. When the code is ready,
run make smoke, which starts two fake upstreams and exercises routing, failover, recovery and
streaming over real HTTP.
Configure
crossbar reads one TOML file. This is example.toml:
# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production
db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
poll_interval = "1s" # 60s in production; 1s makes the smoke run quick
lease_idle = "30m" # a conversation idle this long loses its host
retention = "180d" # per-request rows older than this are rolled up daily
queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that
[hosts.alpha]
base_url = "http://127.0.0.1:18081" # e.g. http://straylight.<tailnet>:11434
weight = 1.0
models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }
[hosts.beta]
base_url = "http://127.0.0.1:18082" # e.g. http://titan.<tailnet>:8081
weight = 2.0
models = { "ornith-1.5-35b-a3b" = { parallel = 2 } }
# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
[routes.opencode-a]
hosts = ["alpha", "beta"]
default_model = "ornith-1.5-35b-a3b"
[routes.hermes-x]
hosts = ["beta", "alpha"]
| Key | Meaning |
|---|---|
listen |
Where crossbar binds. A tailnet address, never 0.0.0.0. |
db |
SQLite file holding leases and the accounting rows. |
lease_idle |
A conversation idle this long loses its host. |
retention |
Per-request rows older than this are rolled up daily. |
poll_interval |
How often each host is health-checked. 60s in production; 1s makes the smoke run quick. |
queue_max |
Waiting places per (host, model) beyond parallel; a full queue returns 503. |
hosts.<name>.base_url |
The llama-server base URL this host serves. |
hosts.<name>.weight |
Relative share of new requests this host receives. |
hosts.<name>.models |
The models this host serves, with per-model parallel tuning. |
routes.<name>.hosts |
Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
routes.<name>.default_model |
Model used when a request omits one; must be served by a host in the route. |
identity |
"off" (default), "tailscale", or "header"; see below. |
hosts.<name>.wake |
A wake-on-LAN target (mac, broadcast, wait) so crossbar can rouse a sleeping host when nothing else can take a new lease. |
routes.<name>.peers |
The tailnet nodes allowed to reach the route, with identity = "tailscale"; see below. |
Run
Copy the binary, the config and the unit into place, reload systemd, and start it:
install -m 0755 bin/crossbar /usr/local/bin/crossbar
install -d -m 0755 /etc/crossbar
install -m 0644 crossbar.toml /etc/crossbar/crossbar.toml
install -m 0644 deploy/crossbar.service /etc/systemd/system/crossbar.service
systemctl daemon-reload
systemctl enable --now crossbar
Point clients at it
OpenCode, one provider for every project. Each instance is launched as
CROSSBAR_ROUTE="$(basename "$PWD")-$$" opencode:
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
"options": { "baseURL": "http://crossbar.<tailnet>:7777/{env:CROSSBAR_ROUTE}/v1" },
"models": { "ornith-1.5-35b-a3b": {} } } }
Hermes, in config.yaml:
custom_providers:
- name: crossbar
base_url: http://crossbar.<tailnet>:7777/hermes-<agent>/v1
models: { ornith-1.5-35b-a3b: {} }
The route name in the URL must exist in [routes]; unknown routes are 404. A client may instead
name the route on an X-Crossbar-Route header and point at the bare /v1 base:
curl -H 'X-Crossbar-Route: opencode-a' \
https://crossbar.<tailnet>:7777/v1/chat/completions
Operate
The operator's API lives under /_crossbar/. Every call returns 200 with a small JSON body unless
stated otherwise.
GET /_crossbar/hosts reports every host's health, loaded models, live concurrency from the
limiter, drain state and the context sizes the poller learned (n_ctx/slots from a single
server's /props, models per loaded model from /props?model=; 0 or absent means unknown):
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}}
On a llama-server router only models whose status.value is "loaded" count as loaded, and
crossbar asks /props?model=X only for those: asking about an unloaded model would make the
router load it.
GET /_crossbar/routes reports each route's candidate hosts, default model, any pin and its live
leases:
{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}
POST /_crossbar/routes/{route} pins a route to a host ({"host":"alpha","pin":true}) or releases
it and clears the pin ({"release":true}):
{"ok":true}
POST /_crossbar/hosts/{host} sets or clears drain ({"drain":true}); a draining host takes no
new conversations but keeps its existing leases:
{"ok":true}
GET /_crossbar/usage summarizes the accounting rows, grouped by by=host, by=model or
by=route (the default). Ask for JSON, or a fixed-width table with Accept: text/plain:
[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
key requests errors busy_ms queued_ms prompt cached completion cache_hit
hermes-x 1 0 1 0 100 90 10 0.90
opencode-a 1 0 3 0 100 90 10 0.90
GET /_crossbar/metrics emits the Prometheus text exposition for request counts, token totals,
queue wait, host health and live slots:
# TYPE crossbar_requests_total counter
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
# TYPE crossbar_host_healthy gauge
crossbar_host_healthy{host="alpha"} 1
crossbar_host_healthy{host="beta"} 1
Context guard
With unified KV a host's usable context per request is its context size divided by its slots.
crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it
with the leased host's per-slot context for that model. A prompt that fits stays put. One that
does not fit is moved to a healthy host on the route where it does fit (the lease moves with
it, so the conversation stays there), and the response carries
X-Crossbar-Ctx: moved:<from> + > + <to> — for example moved:small>big. When no host can
fit it, the answer is 400 {"error":"prompt too large","estimate":<tokens>,"max":<largest per-slot context among hosts that have the model loaded>}. Hosts whose context is unknown are never
blocked by the guard.
Wake
When a route has no healthy host left and at least one candidate lists a wake target, crossbar
sends that host a wake-on-LAN magic packet, in route order, and retries the lease once. A host that
wakes up takes the conversation; if none wakes, the request gets 503 {"error":"no healthy host", "woke":["<hosts tried>"]}. The context-size guard wakes a sleeping host the same way before it
answers 400 prompt too large, when no healthy host's per-slot context can fit the prompt.
Identity
identity gates who may use a route. With the default "off" every request is admitted. With
"tailscale", a route that lists peers answers 403 to any caller whose tailnet address is not
one of them (checked with tailscale whois):
[routes.hermes-x]
hosts = ["beta", "alpha"]
peers = ["talos"]
"header" trusts the X-Crossbar-Peer header instead and needs no tailnet; it is insecure and for
tests only, so crossbar logs a warning when it starts in that mode.
What v2 does not do
Request coalescing, /slots and TLS are out of scope for v2; see PLAN.md.