Files

256 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# crossbar
crossbar is an affinity router in front of several `llama-server` routers. A client's identity is
the first path segment of its base URL — its route. Each conversation takes a sticky lease on one
host, chosen for the most free slots for its model times weight, and streams the answer back
incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all
new in v1.
## Build
Build everything with `make build`; the binaries land in `bin/`. Check the work with `make gate`,
which runs the formatter, vet, tests and line-length check with no network. When the code is ready,
run `make smoke`, which starts two fake upstreams and exercises routing, failover, recovery and
streaming over real HTTP.
## Configure
crossbar reads one TOML file. This is `example.toml`:
```toml
# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production
db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
poll_interval = "1s" # 60s in production; 1s makes the smoke run quick
lease_idle = "30m" # a conversation idle this long loses its host
retention = "180d" # per-request rows older than this are rolled up daily
queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that
[hosts.alpha]
base_url = "http://127.0.0.1:18081" # e.g. http://straylight.<tailnet>:11434
weight = 1.0
models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }
[hosts.beta]
base_url = "http://127.0.0.1:18082" # e.g. http://titan.<tailnet>:8081
weight = 2.0
models = { "ornith-1.5-35b-a3b" = { parallel = 2 } }
# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
[routes.opencode-a]
hosts = ["alpha", "beta"]
default_model = "ornith-1.5-35b-a3b"
[routes.hermes-x]
hosts = ["beta", "alpha"]
```
| Key | Meaning |
| --- | --- |
| `listen` | Where crossbar binds. A tailnet address, never `0.0.0.0`. |
| `db` | SQLite file holding leases and the accounting rows. |
| `lease_idle` | A conversation idle this long loses its host. |
| `retention` | Per-request rows older than this are rolled up daily. |
| `poll_interval` | How often each host is health-checked. 60s in production; 1s makes the smoke run quick. |
| `queue_max` | Waiting places per (host, model) beyond `parallel`; a full queue returns 503. |
| `hosts.<name>.base_url` | The llama-server base URL this host serves. |
| `hosts.<name>.weight` | Relative share of new requests this host receives. |
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
| `routes.<name>.affinity` | `"conversation"` (default, one lease per conversation) or `"route"` (one lease for the whole route); see "Clients that manage their own slots". |
| `routes.<name>.queue` | `false` leaves queueing to the client's own llama-server slot; the default counts requests in crossbar's per-(host, model) queue. |
| `routes.<name>.listen` | A host:port for the route's own listener, every request there is this route; see "Clients that manage their own slots". |
| `identity` | `"off"` (default), `"tailscale"`, or `"header"`; see below. |
| `hosts.<name>.wake` | A wake-on-LAN target (`mac`, `broadcast`, `wait`) so crossbar can rouse a sleeping host when nothing else can take a new lease. |
| `routes.<name>.peers` | The tailnet nodes allowed to reach the route, with `identity = "tailscale"`; see below. |
## Run
Copy the binary, the config and the unit into place, reload systemd, and start it:
```sh
install -m 0755 bin/crossbar /usr/local/bin/crossbar
install -d -m 0755 /etc/crossbar
install -m 0644 crossbar.toml /etc/crossbar/crossbar.toml
install -m 0644 deploy/crossbar.service /etc/systemd/system/crossbar.service
systemctl daemon-reload
systemctl enable --now crossbar
```
## Point clients at it
OpenCode, one provider for every project. Each instance is launched as
`CROSSBAR_ROUTE="$(basename "$PWD")-$$" opencode`:
```jsonc
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
"options": { "baseURL": "http://crossbar.<tailnet>:7777/{env:CROSSBAR_ROUTE}/v1" },
"models": { "ornith-1.5-35b-a3b": {} } } }
```
Hermes, in `config.yaml`:
```yaml
custom_providers:
- name: crossbar
base_url: http://crossbar.<tailnet>:7777/hermes-<agent>/v1
models: { ornith-1.5-35b-a3b: {} }
```
The route name in the URL must exist in `[routes]`; unknown routes are 404. A client may instead
name the route on an `X-Crossbar-Route` header and point at the bare `/v1` base:
```sh
curl -H 'X-Crossbar-Route: opencode-a' \
https://crossbar.<tailnet>:7777/v1/chat/completions
```
## Clients that manage their own slots
Some clients connect to one crossbar address and manage a llama-server slot themselves: they pin
`id_slot`, poll `/slots`, and steer a running completion through
`/v1/chat/completions/control`. Boxmaker's `inferproxy` is one. crossbar serves such a
client from a route that has its own `listen` address and `affinity = "route"`, so the whole route
lives on one host:
```toml
# a client that manages its own llama-server slot (it pins id_slot, polls /slots, steers a
# running completion through /v1/chat/completions/control) and cannot put a route in the path.
# The route gets its own port; every request there is this route and the path goes upstream as is.
[routes.boxmaker-a]
hosts = ["beta", "alpha"]
default_model = "ornith-1.5-35b-a3b"
listen = "127.0.0.1:17801" # a tailnet address in production; never the main listen address
affinity = "route" # one lease for the whole route, not one per conversation
queue = false # counted as load but never held or refused: the server's own slot queue does that
```
Every request to that address is this route, with its whole path passed upstream unchanged (there is
no route segment to strip), so it runs through `Handler.ForRoute` rather than the usual
`/{route}/` path. The address must split into a host and a numeric port, be unique across routes,
not equal the top-level `listen`, and not be on a template route — crossbar refuses any of those at
start-up.
A few things about how crossbar treats those requests:
- **Control calls take no slot.** A GET or HEAD on any allowed path, and a POST to exactly
`/tokenize` or `/v1/chat/completions/control`, is a control call. It follows the route's single
lease but takes no slot, skips the context guard, and writes no accounting row: it is sent beside
its own stream, so it must never wait for or hold a slot. A chat completion on `/v1/chat/completions`
is not a control call.
- **`/slots` and `/tokenize` are proxied; `/slots/<id>` actions are not.** Only the bare `/slots`
path is allowed, so an action on a specific slot id is not forwarded.
- **A GET's model comes from its `?model=` query** (there is no body to read), which is how
`/slots?model=shared` learns which model's slots to report.
- **The admin API is not served on a route listener.** `/_crossbar/hosts` there, and any prefixed
path such as `/boxmaker-a/v1/models`, are 404.
## Operate
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
stated otherwise.
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
limiter, drain state and the context sizes the poller learned (`n_ctx`/`slots` from a single
server's `/props`, `models` per loaded model from `/props?model=`; 0 or absent means unknown):
```json
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}}
```
On a llama-server **router** only models whose `status.value` is `"loaded"` count as loaded, and
crossbar asks `/props?model=X` only for those: asking about an unloaded model would make the
router load it.
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
leases:
```json
{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}
```
`POST /_crossbar/routes/{route}` pins a route to a host (`{"host":"alpha","pin":true}`) or releases
it and clears the pin (`{"release":true}`):
```json
{"ok":true}
```
`POST /_crossbar/hosts/{host}` sets or clears drain (`{"drain":true}`); a draining host takes no
new conversations but keeps its existing leases:
```json
{"ok":true}
```
`GET /_crossbar/usage` summarizes the accounting rows, grouped by `by=host`, `by=model` or
`by=route` (the default). Ask for JSON, or a fixed-width table with `Accept: text/plain`:
```json
[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
```
```
key requests errors busy_ms queued_ms prompt cached completion cache_hit
hermes-x 1 0 1 0 100 90 10 0.90
opencode-a 1 0 3 0 100 90 10 0.90
```
`GET /_crossbar/metrics` emits the Prometheus text exposition for request counts, token totals,
queue wait, host health and live slots:
```
# TYPE crossbar_requests_total counter
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
# TYPE crossbar_host_healthy gauge
crossbar_host_healthy{host="alpha"} 1
crossbar_host_healthy{host="beta"} 1
```
## Context guard
With unified KV a host's usable context per request is its context size divided by its slots.
crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it
with the leased host's per-slot context for that model. A prompt that fits stays put. One that
does not fit is moved to a healthy host on the route where it does fit (the lease moves with
it, so the conversation stays there), and the response carries
`X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can
fit it, the answer is a `400` in llama-server's own overflow shape, so a client that handles the
server's error handles crossbar's refusal too:
```json
{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":<tokens>,"n_ctx":<largest per-slot context among hosts that have the model loaded>}}
```
Hosts whose context is unknown are never blocked by the guard.
## Wake
When a route has no healthy host left and at least one candidate lists a `wake` target, crossbar
sends that host a wake-on-LAN magic packet, in route order, and retries the lease once. A host that
wakes up takes the conversation; if none wakes, the request gets `503 {"error":"no healthy host",
"woke":["<hosts tried>"]}`. The context-size guard wakes a sleeping host the same way before it
answers `400 prompt too large`, when no healthy host's per-slot context can fit the prompt.
## Identity
`identity` gates who may use a route. With the default `"off"` every request is admitted. With
`"tailscale"`, a route that lists `peers` answers `403` to any caller whose tailnet address is not
one of them (checked with `tailscale whois`):
```toml
[routes.hermes-x]
hosts = ["beta", "alpha"]
peers = ["talos"]
```
`"header"` trusts the `X-Crossbar-Peer` header instead and needs no tailnet; it is insecure and for
tests only, so crossbar logs a warning when it starts in that mode.
## What v2 does not do
Request coalescing and TLS are out of scope for v2; see `PLAN.md`.