Smoke run for v1; README for leases, admin and accounting
Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
This commit is contained in:
@@ -1,8 +1,10 @@
|
||||
# crossbar
|
||||
|
||||
crossbar is an affinity router in front of several `llama-server` routers. A client's identity is
|
||||
the first path segment of its base URL; v0 routes each request to the first healthy host on that
|
||||
route's list and streams the answer back unbuffered.
|
||||
the first path segment of its base URL — its route. Each conversation takes a sticky lease on one
|
||||
host, chosen for the most free slots for its model times weight, and streams the answer back
|
||||
incrementally with the usage chunk intact. Pins, drains, queueing, leases and accounting are all
|
||||
new in v1.
|
||||
|
||||
## Build
|
||||
|
||||
@@ -16,22 +18,26 @@ streaming over real HTTP.
|
||||
crossbar reads one TOML file. This is `example.toml`:
|
||||
|
||||
```toml
|
||||
# crossbar example configuration. Replace <tailnet> and the addresses with your own.
|
||||
# crossbar example configuration (v1). Replace <tailnet> and the addresses with your own.
|
||||
listen = "127.0.0.1:17777" # never 0.0.0.0 — bind the tailnet address in production
|
||||
db = "crossbar.db" # SQLite: leases + accounting (WAL). /var/lib/crossbar/crossbar.db under systemd
|
||||
poll_interval = "1s" # 60s in production; 1s makes the smoke run quick
|
||||
queue_max = 8
|
||||
lease_idle = "30m" # a conversation idle this long loses its host
|
||||
retention = "180d" # per-request rows older than this are rolled up daily
|
||||
queue_max = 1 # waiting places per (host, model) beyond `parallel`; 503 past that
|
||||
|
||||
[hosts.alpha]
|
||||
base_url = "http://127.0.0.1:18081" # e.g. http://straylight.<tailnet>:11434
|
||||
weight = 1.0
|
||||
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "small-9b" = { parallel = 6 } }
|
||||
models = { "ornith-1.5-35b-a3b" = { parallel = 1 }, "small-9b" = { parallel = 6 } }
|
||||
|
||||
[hosts.beta]
|
||||
base_url = "http://127.0.0.1:18082" # e.g. http://titan.<tailnet>:8081
|
||||
weight = 2.0
|
||||
models = { "ornith-1.5-35b-a3b" = { parallel = 4 } }
|
||||
models = { "ornith-1.5-35b-a3b" = { parallel = 2 } }
|
||||
|
||||
# v0: a route is a preference list; the first healthy host that has the model wins.
|
||||
# v1: a route is a set of candidate hosts; each conversation gets a sticky lease on the host with
|
||||
# the most free slots × weight at the time it starts. Pins and drains come from the admin API.
|
||||
[routes.opencode-a]
|
||||
hosts = ["alpha", "beta"]
|
||||
default_model = "ornith-1.5-35b-a3b"
|
||||
@@ -43,12 +49,15 @@ hosts = ["beta", "alpha"]
|
||||
| Key | Meaning |
|
||||
| --- | --- |
|
||||
| `listen` | Where crossbar binds. A tailnet address, never `0.0.0.0`. |
|
||||
| `db` | SQLite file holding leases and the accounting rows. |
|
||||
| `lease_idle` | A conversation idle this long loses its host. |
|
||||
| `retention` | Per-request rows older than this are rolled up daily. |
|
||||
| `poll_interval` | How often each host is health-checked. 60s in production; 1s makes the smoke run quick. |
|
||||
| `queue_max` | Reserved for v1 queueing; no effect in v0. |
|
||||
| `queue_max` | Waiting places per (host, model) beyond `parallel`; a full queue returns 503. |
|
||||
| `hosts.<name>.base_url` | The llama-server base URL this host serves. |
|
||||
| `hosts.<name>.weight` | Relative share of new routes this host receives. |
|
||||
| `hosts.<name>.weight` | Relative share of new requests this host receives. |
|
||||
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
|
||||
| `routes.<name>.hosts` | Preference order: the first healthy host that serves the model wins. |
|
||||
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
|
||||
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
|
||||
|
||||
## Run
|
||||
@@ -84,23 +93,73 @@ custom_providers:
|
||||
models: { ornith-1.5-35b-a3b: {} }
|
||||
```
|
||||
|
||||
The route name in the URL must exist in `[routes]`; unknown routes are 404.
|
||||
The route name in the URL must exist in `[routes]`; unknown routes are 404. A client may instead
|
||||
name the route on an `X-Crossbar-Route` header and point at the bare `/v1` base:
|
||||
|
||||
## Inspect
|
||||
|
||||
`GET /_crossbar/hosts` reports every host's health and loaded models:
|
||||
|
||||
```json
|
||||
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T09:34:18Z","last_err":""},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T09:34:18Z","last_err":""}}
|
||||
```sh
|
||||
curl -H 'X-Crossbar-Route: opencode-a' \
|
||||
https://crossbar.<tailnet>:7777/v1/chat/completions
|
||||
```
|
||||
|
||||
`GET /_crossbar/routes` reports each route's preference order and default model:
|
||||
## Operate
|
||||
|
||||
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
|
||||
stated otherwise.
|
||||
|
||||
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
|
||||
limiter and drain state:
|
||||
|
||||
```json
|
||||
{"hermes-x":{"hosts":["beta","alpha"],"default_model":""},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b"}}
|
||||
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false}}
|
||||
```
|
||||
|
||||
## What v0 does not do
|
||||
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
|
||||
leases:
|
||||
|
||||
Leases and stickiness, SQLite, `/slots`, queueing and wake-on-LAN are out of scope for v0; see
|
||||
`PLAN.md`.
|
||||
```json
|
||||
{"hermes-x":{"hosts":["beta","alpha"],"default_model":"","pinned":"","leases":[]},"opencode-a":{"hosts":["alpha","beta"],"default_model":"ornith-1.5-35b-a3b","pinned":"","leases":[]}}
|
||||
```
|
||||
|
||||
`POST /_crossbar/routes/{route}` pins a route to a host (`{"host":"alpha","pin":true}`) or releases
|
||||
it and clears the pin (`{"release":true}`):
|
||||
|
||||
```json
|
||||
{"ok":true}
|
||||
```
|
||||
|
||||
`POST /_crossbar/hosts/{host}` sets or clears drain (`{"drain":true}`); a draining host takes no
|
||||
new conversations but keeps its existing leases:
|
||||
|
||||
```json
|
||||
{"ok":true}
|
||||
```
|
||||
|
||||
`GET /_crossbar/usage` summarizes the accounting rows, grouped by `by=host`, `by=model` or
|
||||
`by=route` (the default). Ask for JSON, or a fixed-width table with `Accept: text/plain`:
|
||||
|
||||
```json
|
||||
[{"key":"beta","requests":2,"errors":0,"busy_ms":4,"queued_ms":0,"prompt_tokens":200,"cached_tokens":180,"completion_tokens":20}]
|
||||
```
|
||||
|
||||
```
|
||||
key requests errors busy_ms queued_ms prompt cached completion cache_hit
|
||||
hermes-x 1 0 1 0 100 90 10 0.90
|
||||
opencode-a 1 0 3 0 100 90 10 0.90
|
||||
```
|
||||
|
||||
`GET /_crossbar/metrics` emits the Prometheus text exposition for request counts, token totals,
|
||||
queue wait, host health and live slots:
|
||||
|
||||
```
|
||||
# TYPE crossbar_requests_total counter
|
||||
crossbar_requests_total{route="hermes-x",host="beta",status="200"} 1
|
||||
crossbar_requests_total{route="opencode-a",host="beta",status="200"} 1
|
||||
# TYPE crossbar_host_healthy gauge
|
||||
crossbar_host_healthy{host="alpha"} 1
|
||||
crossbar_host_healthy{host="beta"} 1
|
||||
```
|
||||
|
||||
## What v1 does not do
|
||||
|
||||
The context-size guard, wake-on-LAN, Tailscale identity and `/slots` are out of scope for v1; see
|
||||
`PLAN.md` v2.
|
||||
|
||||
Reference in New Issue
Block a user