# crossbar crossbar is an affinity router in front of several `llama-server` routers. A client's identity is the first path segment of its base URL — its route. Each conversation takes a sticky lease on one host, chosen for the most free slots for its model times weight, and streams the answer back incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all new in v1. ## Build Build everything with `make build`; the binaries land in `bin/`. Check the work with `make gate`, which runs the formatter, vet, tests and line-length check with no network. When the code is ready, run `make smoke`, which starts two fake upstreams and exercises routing, failover, recovery and streaming over real HTTP. ## Configure crossbar reads one TOML file. This is `example.toml`: ```toml # crossbar example configuration (v1). Replace and the addresses with your own. listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd poll_interval = "1s" # 60s in production; 1s makes the smoke run quick lease_idle = "30m" # a conversation idle this long loses its host retention = "180d" # per-request rows older than this are rolled up daily queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that [hosts.alpha] base_url = "http://127.0.0.1:18081" # e.g. http://straylight.:11434 weight = 1.0 models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } } [hosts.beta] base_url = "http://127.0.0.1:18082" # e.g. http://titan.:8081 weight = 2.0 models = { "ornith-1.5-35b-a3b" = { parallel = 2 } } # v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with # the most free slots × weight at the time it starts. Pins and drains come from the admin API. [routes.opencode-a] hosts = ["alpha", "beta"] default_model = "ornith-1.5-35b-a3b" [routes.hermes-x] hosts = ["beta", "alpha"] ``` | Key | Meaning | | --- | --- | | `listen` | Where crossbar binds. A tailnet address, never `0.0.0.0`. | | `db` | SQLite file holding leases and the accounting rows. | | `lease_idle` | A conversation idle this long loses its host. | | `retention` | Per-request rows older than this are rolled up daily. | | `poll_interval` | How often each host is health-checked. 60s in production; 1s makes the smoke run quick. | | `queue_max` | Waiting places per (host, model) beyond `parallel`; a full queue returns 503. | | `hosts..base_url` | The llama-server base URL this host serves. | | `hosts..weight` | Relative share of new requests this host receives. | | `hosts..models` | The models this host serves, with per-model parallel tuning. | | `routes..hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. | | `routes..default_model` | Model used when a request omits one; must be served by a host in the route. | | `routes..affinity` | `"conversation"` (default, one lease per conversation) or `"route"` (one lease for the whole route); see "Clients that manage their own slots". | | `routes..queue` | `false` leaves queueing to the client's own llama-server slot; the default counts requests in crossbar's per-(host, model) queue. | | `routes..listen` | A host:port for the route's own listener, every request there is this route; see "Clients that manage their own slots". | | `identity` | `"off"` (default), `"tailscale"`, or `"header"`; see below. | | `hosts..wake` | A wake-on-LAN target (`mac`, `broadcast`, `wait`) so crossbar can rouse a sleeping host when nothing else can take a new lease. | | `routes..peers` | The tailnet nodes allowed to reach the route, with `identity = "tailscale"`; see below. | ## Run Copy the binary, the config and the unit into place, reload systemd, and start it: ```sh install -m 0755 bin/crossbar /usr/local/bin/crossbar install -d -m 0755 /etc/crossbar install -m 0644 crossbar.toml /etc/crossbar/crossbar.toml install -m 0644 deploy/crossbar.service /etc/systemd/system/crossbar.service systemctl daemon-reload systemctl enable --now crossbar ``` ## Point clients at it OpenCode, one provider for every project. Each instance is launched as `CROSSBAR_ROUTE="$(basename "$PWD")-$$" opencode`: ```jsonc "provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible", "options": { "baseURL": "http://crossbar.:7777/{env:CROSSBAR_ROUTE}/v1" }, "models": { "ornith-1.5-35b-a3b": {} } } } ``` Hermes, in `config.yaml`: ```yaml custom_providers: - name: crossbar base_url: http://crossbar.:7777/hermes-/v1 models: { ornith-1.5-35b-a3b: {} } ``` The route name in the URL must exist in `[routes]`; unknown routes are 404. A client may instead name the route on an `X-Crossbar-Route` header and point at the bare `/v1` base: ```sh curl -H 'X-Crossbar-Route: opencode-a' \ https://crossbar.:7777/v1/chat/completions ``` ## Clients that manage their own slots Some clients connect to one crossbar address and manage a llama-server slot themselves: they pin `id_slot`, poll `/slots`, and steer a running completion through `/v1/chat/completions/control`. Boxmaker's `inferproxy` is one. crossbar serves such a client from a route that has its own `listen` address and `affinity = "route"`, so the whole route lives on one host: ```toml # a client that manages its own llama-server slot (it pins id_slot, polls /slots, steers a # running completion through /v1/chat/completions/control) and cannot put a route in the path. # The route gets its own port; every request there is this route and the path goes upstream as is. [routes.boxmaker-a] hosts = ["beta", "alpha"] default_model = "ornith-1.5-35b-a3b" listen = "127.0.0.1:17801" # a tailnet address in production; never the main listen address affinity = "route" # one lease for the whole route, not one per conversation queue = false # counted as load but never held or refused: the server's own slot queue does that ``` Every request to that address is this route, with its whole path passed upstream unchanged (there is no route segment to strip), so it runs through `Handler.ForRoute` rather than the usual `/{route}/` path. The address must split into a host and a numeric port, be unique across routes, not equal the top-level `listen`, and not be on a template route — crossbar refuses any of those at start-up. A few things about how crossbar treats those requests: - **Control calls take no slot.** A GET or HEAD on any allowed path, and a POST to exactly `/tokenize` or `/v1/chat/completions/control`, is a control call. It follows the route's single lease but takes no slot, skips the context guard, and writes no accounting row: it is sent beside its own stream, so it must never wait for or hold a slot. A chat completion on `/v1/chat/completions` is not a control call. - **`/slots` and `/tokenize` are proxied; `/slots/` actions are not.** Only the bare `/slots` path is allowed, so an action on a specific slot id is not forwarded. - **A GET's model comes from its `?model=` query** (there is no body to read), which is how `/slots?model=shared` learns which model's slots to report. - **The admin API is not served on a route listener.** `/_crossbar/hosts` there, and any prefixed path such as `/boxmaker-a/v1/models`, are 404. ## Operate The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless stated otherwise. `GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the limiter, drain state and the context sizes the poller learned (`n_ctx`/`slots` from a single server's `/props`, `models` per loaded model from `/props?model=`; 0 or absent means unknown): ```json {"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}} ``` On a llama-server **router** only models whose `status.value` is `"loaded"` count as loaded, and crossbar asks `/props?model=X` only for those: asking about an unloaded model would make the router load it. `GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live leases: ```json {"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}} ``` `POST /_crossbar/routes/{route}` pins a route to a host (`{"host":"alpha","pin":true}`) or releases it and clears the pin (`{"release":true}`): ```json {"ok":true} ``` `POST /_crossbar/hosts/{host}` sets or clears drain (`{"drain":true}`); a draining host takes no new conversations but keeps its existing leases: ```json {"ok":true} ``` `GET /_crossbar/usage` summarizes the accounting rows, grouped by `by=host`, `by=model` or `by=route` (the default). Ask for JSON, or a fixed-width table with `Accept: text/plain`: ```json [{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}] ``` ``` key requests errors busy_ms queued_ms prompt cached completion cache_hit hermes-x 1 0 1 0 100 90 10 0.90 opencode-a 1 0 3 0 100 90 10 0.90 ``` `GET /_crossbar/metrics` emits the Prometheus text exposition for request counts, token totals, queue wait, host health and live slots: ``` # TYPE crossbar_requests_total counter crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1 crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1 # TYPE crossbar_host_healthy gauge crossbar_host_healthy{host="alpha"} 1 crossbar_host_healthy{host="beta"} 1 ``` ## Context guard With unified KV a host's usable context per request is its context size divided by its slots. crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it with the leased host's per-slot context for that model. A prompt that fits stays put. One that does not fit is moved to a healthy host on the route where it does fit (the lease moves with it, so the conversation stays there), and the response carries `X-Crossbar-Ctx: moved:` + `>` + `` — for example `moved:small>big`. When no host can fit it, the answer is a `400` in llama-server's own overflow shape, so a client that handles the server's error handles crossbar's refusal too: ```json {"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":,"n_ctx":}} ``` Hosts whose context is unknown are never blocked by the guard. ## Wake When a route has no healthy host left and at least one candidate lists a `wake` target, crossbar sends that host a wake-on-LAN magic packet, in route order, and retries the lease once. A host that wakes up takes the conversation; if none wakes, the request gets `503 {"error":"no healthy host", "woke":[""]}`. The context-size guard wakes a sleeping host the same way before it answers `400 prompt too large`, when no healthy host's per-slot context can fit the prompt. ## Identity `identity` gates who may use a route. With the default `"off"` every request is admitted. With `"tailscale"`, a route that lists `peers` answers `403` to any caller whose tailnet address is not one of them (checked with `tailscale whois`): ```toml [routes.hermes-x] hosts = ["beta", "alpha"] peers = ["talos"] ``` `"header"` trusts the `X-Crossbar-Peer` header instead and needs no tailnet; it is insecure and for tests only, so crossbar logs a warning when it starts in that mode. ## What v2 does not do Request coalescing and TLS are out of scope for v2; see `PLAN.md`.