Smoke run for v1; README for leases, admin and accounting

Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
This commit is contained in:
2026-09-25 06:56:54 -07:00
parent 587ec7a1ec
commit cf2aa24393
5 changed files with 178 additions and 55 deletions
+81 -22
View File
@@ -1,8 +1,10 @@
# crossbar
crossbar is an affinity router in front of several `llama-server` routers. A client's identity is
the first path segment of its base URL; v0 routes each request to the first healthy host on that
route's list and streams the answer back unbuffered.
the first path segment of its base URL — its route. Each conversation takes a sticky lease on one
host, chosen for the most free slots for its model times weight, and streams the answer back
incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all
new in v1.
## Build
@@ -16,22 +18,26 @@ streaming over real HTTP.
crossbar reads one TOML file. This is `example.toml`:
```toml
# crossbar example configuration. Replace <tailnet> and the addresses with your own.
# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production
db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
poll_interval = "1s" # 60s in production; 1s makes the smoke run quick
queue_max = 8
lease_idle = "30m" # a conversation idle this long loses its host
retention = "180d" # per-request rows older than this are rolled up daily
queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that
[hosts.alpha]
base_url = "http://127.0.0.1:18081" # e.g. http://straylight.<tailnet>:11434
weight = 1.0
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "small-9b" = { parallel = 6 } }
models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }
[hosts.beta]
base_url = "http://127.0.0.1:18082" # e.g. http://titan.<tailnet>:8081
weight = 2.0
models = { "ornith-1.5-35b-a3b" = { parallel = 4 } }
models = { "ornith-1.5-35b-a3b" = { parallel = 2 } }
# v0: a route is a preference list; the first healthy host that has the model wins.
# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
[routes.opencode-a]
hosts = ["alpha", "beta"]
default_model = "ornith-1.5-35b-a3b"
@@ -43,12 +49,15 @@ hosts = ["beta", "alpha"]
| Key | Meaning |
| --- | --- |
| `listen` | Where crossbar binds. A tailnet address, never `0.0.0.0`. |
| `db` | SQLite file holding leases and the accounting rows. |
| `lease_idle` | A conversation idle this long loses its host. |
| `retention` | Per-request rows older than this are rolled up daily. |
| `poll_interval` | How often each host is health-checked. 60s in production; 1s makes the smoke run quick. |
| `queue_max` | Reserved for v1 queueing; no effect in v0. |
| `queue_max` | Waiting places per (host, model) beyond `parallel`; a full queue returns 503. |
| `hosts.<name>.base_url` | The llama-server base URL this host serves. |
| `hosts.<name>.weight` | Relative share of new routes this host receives. |
| `hosts.<name>.weight` | Relative share of new requests this host receives. |
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
| `routes.<name>.hosts` | Preference order: the first healthy host that serves the model wins. |
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
## Run
@@ -84,23 +93,73 @@ custom_providers:
models: { ornith-1.5-35b-a3b: {} }
```
The route name in the URL must exist in `[routes]`; unknown routes are 404.
The route name in the URL must exist in `[routes]`; unknown routes are 404. A client may instead
name the route on an `X-Crossbar-Route` header and point at the bare `/v1` base:
## Inspect
`GET /_crossbar/hosts` reports every host's health and loaded models:
```json
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T09:34:18Z","last_err":""},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T09:34:18Z","last_err":""}}
```sh
curl -H 'X-Crossbar-Route: opencode-a' \
https://crossbar.<tailnet>:7777/v1/chat/completions
```
`GET /_crossbar/routes` reports each route's preference order and default model:
## Operate
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
stated otherwise.
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
limiter and drain state:
```json
{"hermes-x":{"hosts":["beta","alpha"],"default_model":""},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b"}}
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false}}
```
## What v0 does not do
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
leases:
Leases and stickiness, SQLite, `/slots`, queueing and wake-on-LAN are out of scope for v0; see
`PLAN.md`.
```json
{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}
```
`POST /_crossbar/routes/{route}` pins a route to a host (`{"host":"alpha","pin":true}`) or releases
it and clears the pin (`{"release":true}`):
```json
{"ok":true}
```
`POST /_crossbar/hosts/{host}` sets or clears drain (`{"drain":true}`); a draining host takes no
new conversations but keeps its existing leases:
```json
{"ok":true}
```
`GET /_crossbar/usage` summarizes the accounting rows, grouped by `by=host`, `by=model` or
`by=route` (the default). Ask for JSON, or a fixed-width table with `Accept: text/plain`:
```json
[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
```
```
key requests errors busy_ms queued_ms prompt cached completion cache_hit
hermes-x 1 0 1 0 100 90 10 0.90
opencode-a 1 0 3 0 100 90 10 0.90
```
`GET /_crossbar/metrics` emits the Prometheus text exposition for request counts, token totals,
queue wait, host health and live slots:
```
# TYPE crossbar_requests_total counter
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
# TYPE crossbar_host_healthy gauge
crossbar_host_healthy{host="alpha"} 1
crossbar_host_healthy{host="beta"} 1
```
## What v1 does not do
The context-size guard, wake-on-LAN, Tailscale identity and `/slots` are out of scope for v1; see
`PLAN.md` v2.