kyle 33fa61bedb Control-plane requests follow the lease but take no slot and write no row
Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
2026-09-25 17:44:00 -07:00

crossbar

crossbar is an affinity router in front of several llama-server routers. A client's identity is the first path segment of its base URL — its route. Each conversation takes a sticky lease on one host, chosen for the most free slots for its model times weight, and streams the answer back incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all new in v1.

Build

Build everything with make build; the binaries land in bin/. Check the work with make gate, which runs the formatter, vet, tests and line-length check with no network. When the code is ready, run make smoke, which starts two fake upstreams and exercises routing, failover, recovery and streaming over real HTTP.

Configure

crossbar reads one TOML file. This is example.toml:

# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
listen        = "127.0.0.1:17777"   # never 0.0.0.0 — bind the tailnet address in production
db            = "crossbar.db"       # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
poll_interval = "1s"                # 60s in production; 1s makes the smoke run quick
lease_idle    = "30m"               # a conversation idle this long loses its host
retention     = "180d"              # per-request rows older than this are rolled up daily
queue_max     = 1                   # waiting places per (host, model) beyond `parallel`; 503 past that

[hosts.alpha]
base_url = "http://127.0.0.1:18081"    # e.g. http://straylight.<tailnet>:11434
weight   = 1.0
models   = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }

[hosts.beta]
base_url = "http://127.0.0.1:18082"    # e.g. http://titan.<tailnet>:8081
weight   = 2.0
models   = { "ornith-1.5-35b-a3b" = { parallel = 2 } }

# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
[routes.opencode-a]
hosts         = ["alpha", "beta"]
default_model = "ornith-1.5-35b-a3b"

[routes.hermes-x]
hosts = ["beta", "alpha"]
Key Meaning
listen Where crossbar binds. A tailnet address, never 0.0.0.0.
db SQLite file holding leases and the accounting rows.
lease_idle A conversation idle this long loses its host.
retention Per-request rows older than this are rolled up daily.
poll_interval How often each host is health-checked. 60s in production; 1s makes the smoke run quick.
queue_max Waiting places per (host, model) beyond parallel; a full queue returns 503.
hosts.<name>.base_url The llama-server base URL this host serves.
hosts.<name>.weight Relative share of new requests this host receives.
hosts.<name>.models The models this host serves, with per-model parallel tuning.
routes.<name>.hosts Candidate hosts, tried in order until one is healthy; a conversation leases one of them.
routes.<name>.default_model Model used when a request omits one; must be served by a host in the route.
identity "off" (default), "tailscale", or "header"; see below.
hosts.<name>.wake A wake-on-LAN target (mac, broadcast, wait) so crossbar can rouse a sleeping host when nothing else can take a new lease.
routes.<name>.peers The tailnet nodes allowed to reach the route, with identity = "tailscale"; see below.

Run

Copy the binary, the config and the unit into place, reload systemd, and start it:

install -m 0755 bin/crossbar /usr/local/bin/crossbar
install -d -m 0755 /etc/crossbar
install -m 0644 crossbar.toml /etc/crossbar/crossbar.toml
install -m 0644 deploy/crossbar.service /etc/systemd/system/crossbar.service
systemctl daemon-reload
systemctl enable --now crossbar

Point clients at it

OpenCode, one provider for every project. Each instance is launched as CROSSBAR_ROUTE="$(basename "$PWD")-$$" opencode:

"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
  "options": { "baseURL": "http://crossbar.<tailnet>:7777/{env:CROSSBAR_ROUTE}/v1" },
  "models": { "ornith-1.5-35b-a3b": {} } } }

Hermes, in config.yaml:

custom_providers:
  - name: crossbar
    base_url: http://crossbar.<tailnet>:7777/hermes-<agent>/v1
    models: { ornith-1.5-35b-a3b: {} }

The route name in the URL must exist in [routes]; unknown routes are 404. A client may instead name the route on an X-Crossbar-Route header and point at the bare /v1 base:

curl -H 'X-Crossbar-Route: opencode-a' \
  https://crossbar.<tailnet>:7777/v1/chat/completions

Operate

The operator's API lives under /_crossbar/. Every call returns 200 with a small JSON body unless stated otherwise.

GET /_crossbar/hosts reports every host's health, loaded models, live concurrency from the limiter, drain state and the context sizes the poller learned (n_ctx/slots from a single server's /props, models per loaded model from /props?model=; 0 or absent means unknown):

{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}}

On a llama-server router only models whose status.value is "loaded" count as loaded, and crossbar asks /props?model=X only for those: asking about an unloaded model would make the router load it.

GET /_crossbar/routes reports each route's candidate hosts, default model, any pin and its live leases:

{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}

POST /_crossbar/routes/{route} pins a route to a host ({"host":"alpha","pin":true}) or releases it and clears the pin ({"release":true}):

{"ok":true}

POST /_crossbar/hosts/{host} sets or clears drain ({"drain":true}); a draining host takes no new conversations but keeps its existing leases:

{"ok":true}

GET /_crossbar/usage summarizes the accounting rows, grouped by by=host, by=model or by=route (the default). Ask for JSON, or a fixed-width table with Accept: text/plain:

[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
key        requests errors busy_ms queued_ms prompt cached completion cache_hit
hermes-x   1        0      1       0         100    90     10         0.90
opencode-a 1        0      3       0         100    90     10         0.90

GET /_crossbar/metrics emits the Prometheus text exposition for request counts, token totals, queue wait, host health and live slots:

# TYPE crossbar_requests_total counter
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
# TYPE crossbar_host_healthy gauge
crossbar_host_healthy{host="alpha"} 1
crossbar_host_healthy{host="beta"} 1

Context guard

With unified KV a host's usable context per request is its context size divided by its slots. crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it with the leased host's per-slot context for that model. A prompt that fits stays put. One that does not fit is moved to a healthy host on the route where it does fit (the lease moves with it, so the conversation stays there), and the response carries X-Crossbar-Ctx: moved:<from> + > + <to> — for example moved:small>big. When no host can fit it, the answer is 400 {"error":"prompt too large","estimate":<tokens>,"max":<largest per-slot context among hosts that have the model loaded>}. Hosts whose context is unknown are never blocked by the guard.

Wake

When a route has no healthy host left and at least one candidate lists a wake target, crossbar sends that host a wake-on-LAN magic packet, in route order, and retries the lease once. A host that wakes up takes the conversation; if none wakes, the request gets 503 {"error":"no healthy host", "woke":["<hosts tried>"]}. The context-size guard wakes a sleeping host the same way before it answers 400 prompt too large, when no healthy host's per-slot context can fit the prompt.

Identity

identity gates who may use a route. With the default "off" every request is admitted. With "tailscale", a route that lists peers answers 403 to any caller whose tailnet address is not one of them (checked with tailscale whois):

[routes.hermes-x]
hosts = ["beta", "alpha"]
peers = ["talos"]

"header" trusts the X-Crossbar-Peer header instead and needs no tailnet; it is insecure and for tests only, so crossbar logs a warning when it starts in that mode.

What v2 does not do

Request coalescing, /slots and TLS are out of scope for v2; see PLAN.md.

S
Description
No description provided
Readme
823 KiB
Languages
Go 96.4%
Shell 3.4%
Makefile 0.2%