166 lines
6.6 KiB
Markdown
166 lines
6.6 KiB
Markdown
# crossbar
|
||
|
||
crossbar is an affinity router in front of several `llama-server` routers. A client's identity is
|
||
the first path segment of its base URL — its route. Each conversation takes a sticky lease on one
|
||
host, chosen for the most free slots for its model times weight, and streams the answer back
|
||
incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all
|
||
new in v1.
|
||
|
||
## Build
|
||
|
||
Build everything with `make build`; the binaries land in `bin/`. Check the work with `make gate`,
|
||
which runs the formatter, vet, tests and line-length check with no network. When the code is ready,
|
||
run `make smoke`, which starts two fake upstreams and exercises routing, failover, recovery and
|
||
streaming over real HTTP.
|
||
|
||
## Configure
|
||
|
||
crossbar reads one TOML file. This is `example.toml`:
|
||
|
||
```toml
|
||
# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
|
||
listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production
|
||
db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
|
||
poll_interval = "1s" # 60s in production; 1s makes the smoke run quick
|
||
lease_idle = "30m" # a conversation idle this long loses its host
|
||
retention = "180d" # per-request rows older than this are rolled up daily
|
||
queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that
|
||
|
||
[hosts.alpha]
|
||
base_url = "http://127.0.0.1:18081" # e.g. http://straylight.<tailnet>:11434
|
||
weight = 1.0
|
||
models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }
|
||
|
||
[hosts.beta]
|
||
base_url = "http://127.0.0.1:18082" # e.g. http://titan.<tailnet>:8081
|
||
weight = 2.0
|
||
models = { "ornith-1.5-35b-a3b" = { parallel = 2 } }
|
||
|
||
# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
|
||
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
|
||
[routes.opencode-a]
|
||
hosts = ["alpha", "beta"]
|
||
default_model = "ornith-1.5-35b-a3b"
|
||
|
||
[routes.hermes-x]
|
||
hosts = ["beta", "alpha"]
|
||
```
|
||
|
||
| Key | Meaning |
|
||
| --- | --- |
|
||
| `listen` | Where crossbar binds. A tailnet address, never `0.0.0.0`. |
|
||
| `db` | SQLite file holding leases and the accounting rows. |
|
||
| `lease_idle` | A conversation idle this long loses its host. |
|
||
| `retention` | Per-request rows older than this are rolled up daily. |
|
||
| `poll_interval` | How often each host is health-checked. 60s in production; 1s makes the smoke run quick. |
|
||
| `queue_max` | Waiting places per (host, model) beyond `parallel`; a full queue returns 503. |
|
||
| `hosts.<name>.base_url` | The llama-server base URL this host serves. |
|
||
| `hosts.<name>.weight` | Relative share of new requests this host receives. |
|
||
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
|
||
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
|
||
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
|
||
|
||
## Run
|
||
|
||
Copy the binary, the config and the unit into place, reload systemd, and start it:
|
||
|
||
```sh
|
||
install -m 0755 bin/crossbar /usr/local/bin/crossbar
|
||
install -d -m 0755 /etc/crossbar
|
||
install -m 0644 crossbar.toml /etc/crossbar/crossbar.toml
|
||
install -m 0644 deploy/crossbar.service /etc/systemd/system/crossbar.service
|
||
systemctl daemon-reload
|
||
systemctl enable --now crossbar
|
||
```
|
||
|
||
## Point clients at it
|
||
|
||
OpenCode, one provider for every project. Each instance is launched as
|
||
`CROSSBAR_ROUTE="$(basename "$PWD")-$$" opencode`:
|
||
|
||
```jsonc
|
||
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
|
||
"options": { "baseURL": "http://crossbar.<tailnet>:7777/{env:CROSSBAR_ROUTE}/v1" },
|
||
"models": { "ornith-1.5-35b-a3b": {} } } }
|
||
```
|
||
|
||
Hermes, in `config.yaml`:
|
||
|
||
```yaml
|
||
custom_providers:
|
||
- name: crossbar
|
||
base_url: http://crossbar.<tailnet>:7777/hermes-<agent>/v1
|
||
models: { ornith-1.5-35b-a3b: {} }
|
||
```
|
||
|
||
The route name in the URL must exist in `[routes]`; unknown routes are 404. A client may instead
|
||
name the route on an `X-Crossbar-Route` header and point at the bare `/v1` base:
|
||
|
||
```sh
|
||
curl -H 'X-Crossbar-Route: opencode-a' \
|
||
https://crossbar.<tailnet>:7777/v1/chat/completions
|
||
```
|
||
|
||
## Operate
|
||
|
||
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
|
||
stated otherwise.
|
||
|
||
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
|
||
limiter and drain state:
|
||
|
||
```json
|
||
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false}}
|
||
```
|
||
|
||
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
|
||
leases:
|
||
|
||
```json
|
||
{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}
|
||
```
|
||
|
||
`POST /_crossbar/routes/{route}` pins a route to a host (`{"host":"alpha","pin":true}`) or releases
|
||
it and clears the pin (`{"release":true}`):
|
||
|
||
```json
|
||
{"ok":true}
|
||
```
|
||
|
||
`POST /_crossbar/hosts/{host}` sets or clears drain (`{"drain":true}`); a draining host takes no
|
||
new conversations but keeps its existing leases:
|
||
|
||
```json
|
||
{"ok":true}
|
||
```
|
||
|
||
`GET /_crossbar/usage` summarizes the accounting rows, grouped by `by=host`, `by=model` or
|
||
`by=route` (the default). Ask for JSON, or a fixed-width table with `Accept: text/plain`:
|
||
|
||
```json
|
||
[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
|
||
```
|
||
|
||
```
|
||
key requests errors busy_ms queued_ms prompt cached completion cache_hit
|
||
hermes-x 1 0 1 0 100 90 10 0.90
|
||
opencode-a 1 0 3 0 100 90 10 0.90
|
||
```
|
||
|
||
`GET /_crossbar/metrics` emits the Prometheus text exposition for request counts, token totals,
|
||
queue wait, host health and live slots:
|
||
|
||
```
|
||
# TYPE crossbar_requests_total counter
|
||
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
|
||
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
|
||
# TYPE crossbar_host_healthy gauge
|
||
crossbar_host_healthy{host="alpha"} 1
|
||
crossbar_host_healthy{host="beta"} 1
|
||
```
|
||
|
||
## What v1 does not do
|
||
|
||
The context-size guard, wake-on-LAN, Tailscale identity and `/slots` are out of scope for v1; see
|
||
`PLAN.md` v2.
|