256 lines
12 KiB
Markdown
256 lines
12 KiB
Markdown
# crossbar
|
||
|
||
crossbar is an affinity router in front of several `llama-server` routers. A client's identity is
|
||
the first path segment of its base URL — its route. Each conversation takes a sticky lease on one
|
||
host, chosen for the most free slots for its model times weight, and streams the answer back
|
||
incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all
|
||
new in v1.
|
||
|
||
## Build
|
||
|
||
Build everything with `make build`; the binaries land in `bin/`. Check the work with `make gate`,
|
||
which runs the formatter, vet, tests and line-length check with no network. When the code is ready,
|
||
run `make smoke`, which starts two fake upstreams and exercises routing, failover, recovery and
|
||
streaming over real HTTP.
|
||
|
||
## Configure
|
||
|
||
crossbar reads one TOML file. This is `example.toml`:
|
||
|
||
```toml
|
||
# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
|
||
listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production
|
||
db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
|
||
poll_interval = "1s" # 60s in production; 1s makes the smoke run quick
|
||
lease_idle = "30m" # a conversation idle this long loses its host
|
||
retention = "180d" # per-request rows older than this are rolled up daily
|
||
queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that
|
||
|
||
[hosts.alpha]
|
||
base_url = "http://127.0.0.1:18081" # e.g. http://straylight.<tailnet>:11434
|
||
weight = 1.0
|
||
models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }
|
||
|
||
[hosts.beta]
|
||
base_url = "http://127.0.0.1:18082" # e.g. http://titan.<tailnet>:8081
|
||
weight = 2.0
|
||
models = { "ornith-1.5-35b-a3b" = { parallel = 2 } }
|
||
|
||
# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
|
||
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
|
||
[routes.opencode-a]
|
||
hosts = ["alpha", "beta"]
|
||
default_model = "ornith-1.5-35b-a3b"
|
||
|
||
[routes.hermes-x]
|
||
hosts = ["beta", "alpha"]
|
||
```
|
||
|
||
| Key | Meaning |
|
||
| --- | --- |
|
||
| `listen` | Where crossbar binds. A tailnet address, never `0.0.0.0`. |
|
||
| `db` | SQLite file holding leases and the accounting rows. |
|
||
| `lease_idle` | A conversation idle this long loses its host. |
|
||
| `retention` | Per-request rows older than this are rolled up daily. |
|
||
| `poll_interval` | How often each host is health-checked. 60s in production; 1s makes the smoke run quick. |
|
||
| `queue_max` | Waiting places per (host, model) beyond `parallel`; a full queue returns 503. |
|
||
| `hosts.<name>.base_url` | The llama-server base URL this host serves. |
|
||
| `hosts.<name>.weight` | Relative share of new requests this host receives. |
|
||
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
|
||
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
|
||
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
|
||
| `routes.<name>.affinity` | `"conversation"` (default, one lease per conversation) or `"route"` (one lease for the whole route); see "Clients that manage their own slots". |
|
||
| `routes.<name>.queue` | `false` leaves queueing to the client's own llama-server slot; the default counts requests in crossbar's per-(host, model) queue. |
|
||
| `routes.<name>.listen` | A host:port for the route's own listener, every request there is this route; see "Clients that manage their own slots". |
|
||
| `identity` | `"off"` (default), `"tailscale"`, or `"header"`; see below. |
|
||
| `hosts.<name>.wake` | A wake-on-LAN target (`mac`, `broadcast`, `wait`) so crossbar can rouse a sleeping host when nothing else can take a new lease. |
|
||
| `routes.<name>.peers` | The tailnet nodes allowed to reach the route, with `identity = "tailscale"`; see below. |
|
||
|
||
## Run
|
||
|
||
Copy the binary, the config and the unit into place, reload systemd, and start it:
|
||
|
||
```sh
|
||
install -m 0755 bin/crossbar /usr/local/bin/crossbar
|
||
install -d -m 0755 /etc/crossbar
|
||
install -m 0644 crossbar.toml /etc/crossbar/crossbar.toml
|
||
install -m 0644 deploy/crossbar.service /etc/systemd/system/crossbar.service
|
||
systemctl daemon-reload
|
||
systemctl enable --now crossbar
|
||
```
|
||
|
||
## Point clients at it
|
||
|
||
OpenCode, one provider for every project. Each instance is launched as
|
||
`CROSSBAR_ROUTE="$(basename "$PWD")-$$" opencode`:
|
||
|
||
```jsonc
|
||
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
|
||
"options": { "baseURL": "http://crossbar.<tailnet>:7777/{env:CROSSBAR_ROUTE}/v1" },
|
||
"models": { "ornith-1.5-35b-a3b": {} } } }
|
||
```
|
||
|
||
Hermes, in `config.yaml`:
|
||
|
||
```yaml
|
||
custom_providers:
|
||
- name: crossbar
|
||
base_url: http://crossbar.<tailnet>:7777/hermes-<agent>/v1
|
||
models: { ornith-1.5-35b-a3b: {} }
|
||
```
|
||
|
||
The route name in the URL must exist in `[routes]`; unknown routes are 404. A client may instead
|
||
name the route on an `X-Crossbar-Route` header and point at the bare `/v1` base:
|
||
|
||
```sh
|
||
curl -H 'X-Crossbar-Route: opencode-a' \
|
||
https://crossbar.<tailnet>:7777/v1/chat/completions
|
||
```
|
||
|
||
## Clients that manage their own slots
|
||
|
||
Some clients connect to one crossbar address and manage a llama-server slot themselves: they pin
|
||
`id_slot`, poll `/slots`, and steer a running completion through
|
||
`/v1/chat/completions/control`. Boxmaker's `inferproxy` is one. crossbar serves such a
|
||
client from a route that has its own `listen` address and `affinity = "route"`, so the whole route
|
||
lives on one host:
|
||
|
||
```toml
|
||
# a client that manages its own llama-server slot (it pins id_slot, polls /slots, steers a
|
||
# running completion through /v1/chat/completions/control) and cannot put a route in the path.
|
||
# The route gets its own port; every request there is this route and the path goes upstream as is.
|
||
[routes.boxmaker-a]
|
||
hosts = ["beta", "alpha"]
|
||
default_model = "ornith-1.5-35b-a3b"
|
||
listen = "127.0.0.1:17801" # a tailnet address in production; never the main listen address
|
||
affinity = "route" # one lease for the whole route, not one per conversation
|
||
queue = false # counted as load but never held or refused: the server's own slot queue does that
|
||
```
|
||
|
||
Every request to that address is this route, with its whole path passed upstream unchanged (there is
|
||
no route segment to strip), so it runs through `Handler.ForRoute` rather than the usual
|
||
`/{route}/` path. The address must split into a host and a numeric port, be unique across routes,
|
||
not equal the top-level `listen`, and not be on a template route — crossbar refuses any of those at
|
||
start-up.
|
||
|
||
A few things about how crossbar treats those requests:
|
||
|
||
- **Control calls take no slot.** A GET or HEAD on any allowed path, and a POST to exactly
|
||
`/tokenize` or `/v1/chat/completions/control`, is a control call. It follows the route's single
|
||
lease but takes no slot, skips the context guard, and writes no accounting row: it is sent beside
|
||
its own stream, so it must never wait for or hold a slot. A chat completion on `/v1/chat/completions`
|
||
is not a control call.
|
||
- **`/slots` and `/tokenize` are proxied; `/slots/<id>` actions are not.** Only the bare `/slots`
|
||
path is allowed, so an action on a specific slot id is not forwarded.
|
||
- **A GET's model comes from its `?model=` query** (there is no body to read), which is how
|
||
`/slots?model=shared` learns which model's slots to report.
|
||
- **The admin API is not served on a route listener.** `/_crossbar/hosts` there, and any prefixed
|
||
path such as `/boxmaker-a/v1/models`, are 404.
|
||
|
||
## Operate
|
||
|
||
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
|
||
stated otherwise.
|
||
|
||
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
|
||
limiter, drain state and the context sizes the poller learned (`n_ctx`/`slots` from a single
|
||
server's `/props`, `models` per loaded model from `/props?model=`; 0 or absent means unknown):
|
||
|
||
```json
|
||
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}}
|
||
```
|
||
|
||
On a llama-server **router** only models whose `status.value` is `"loaded"` count as loaded, and
|
||
crossbar asks `/props?model=X` only for those: asking about an unloaded model would make the
|
||
router load it.
|
||
|
||
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
|
||
leases:
|
||
|
||
```json
|
||
{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}
|
||
```
|
||
|
||
`POST /_crossbar/routes/{route}` pins a route to a host (`{"host":"alpha","pin":true}`) or releases
|
||
it and clears the pin (`{"release":true}`):
|
||
|
||
```json
|
||
{"ok":true}
|
||
```
|
||
|
||
`POST /_crossbar/hosts/{host}` sets or clears drain (`{"drain":true}`); a draining host takes no
|
||
new conversations but keeps its existing leases:
|
||
|
||
```json
|
||
{"ok":true}
|
||
```
|
||
|
||
`GET /_crossbar/usage` summarizes the accounting rows, grouped by `by=host`, `by=model` or
|
||
`by=route` (the default). Ask for JSON, or a fixed-width table with `Accept: text/plain`:
|
||
|
||
```json
|
||
[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
|
||
```
|
||
|
||
```
|
||
key requests errors busy_ms queued_ms prompt cached completion cache_hit
|
||
hermes-x 1 0 1 0 100 90 10 0.90
|
||
opencode-a 1 0 3 0 100 90 10 0.90
|
||
```
|
||
|
||
`GET /_crossbar/metrics` emits the Prometheus text exposition for request counts, token totals,
|
||
queue wait, host health and live slots:
|
||
|
||
```
|
||
# TYPE crossbar_requests_total counter
|
||
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
|
||
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
|
||
# TYPE crossbar_host_healthy gauge
|
||
crossbar_host_healthy{host="alpha"} 1
|
||
crossbar_host_healthy{host="beta"} 1
|
||
```
|
||
|
||
## Context guard
|
||
|
||
With unified KV a host's usable context per request is its context size divided by its slots.
|
||
crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it
|
||
with the leased host's per-slot context for that model. A prompt that fits stays put. One that
|
||
does not fit is moved to a healthy host on the route where it does fit (the lease moves with
|
||
it, so the conversation stays there), and the response carries
|
||
`X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can
|
||
fit it, the answer is a `400` in llama-server's own overflow shape, so a client that handles the
|
||
server's error handles crossbar's refusal too:
|
||
|
||
```json
|
||
{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":<tokens>,"n_ctx":<largest per-slot context among hosts that have the model loaded>}}
|
||
```
|
||
|
||
Hosts whose context is unknown are never blocked by the guard.
|
||
|
||
## Wake
|
||
|
||
When a route has no healthy host left and at least one candidate lists a `wake` target, crossbar
|
||
sends that host a wake-on-LAN magic packet, in route order, and retries the lease once. A host that
|
||
wakes up takes the conversation; if none wakes, the request gets `503 {"error":"no healthy host",
|
||
"woke":["<hosts tried>"]}`. The context-size guard wakes a sleeping host the same way before it
|
||
answers `400 prompt too large`, when no healthy host's per-slot context can fit the prompt.
|
||
|
||
## Identity
|
||
|
||
`identity` gates who may use a route. With the default `"off"` every request is admitted. With
|
||
`"tailscale"`, a route that lists `peers` answers `403` to any caller whose tailnet address is not
|
||
one of them (checked with `tailscale whois`):
|
||
|
||
```toml
|
||
[routes.hermes-x]
|
||
hosts = ["beta", "alpha"]
|
||
peers = ["talos"]
|
||
```
|
||
|
||
`"header"` trusts the `X-Crossbar-Peer` header instead and needs no tailnet; it is insecure and for
|
||
tests only, so crossbar logs a warning when it starts in that mode.
|
||
|
||
## What v2 does not do
|
||
|
||
Request coalescing and TLS are out of scope for v2; see `PLAN.md`.
|