PLAN rev 2: SQLite as state store; per-route/model/host accounting

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 00:44:11 -07:00
co-authored by Claude Fable 5.1
parent 2c6bd7b2ce
commit 9f41ce7935
+54 -4
View File
@@ -1,6 +1,6 @@
# crossbar — an affinity router for the fleet's llama-servers # crossbar — an affinity router for the fleet's llama-servers
**Status:** plan, 2026-09-25. Nothing built. Go. **Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting, Kyle). Nothing built. Go.
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure). **Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
## 1. Problem ## 1. Problem
@@ -71,8 +71,8 @@ no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
lease — the re-prefill is the cost we are avoiding. lease — the re-prefill is the cost we are avoiding.
- **Pinned** leases never move automatically ("project A goes to titan right - **Pinned** leases never move automatically ("project A goes to titan right
now"). **Draining** hosts accept no new leases; existing ones finish. now"). **Draining** hosts accept no new leases; existing ones finish.
- Table persisted to `state.json` after every change; loaded at start so a - Table lives in memory and is written through to SQLite (§7a) on every
crossbar restart does not reshuffle sessions. change; loaded at start so a crossbar restart does not reshuffle sessions.
## 6. Health ## 6. Health
@@ -95,14 +95,62 @@ no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin; - `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
`{ "release": true }` — drop the lease (next request re-chooses). `{ "release": true }` — drop the lease (next request re-chooses).
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`. - `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
- `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups
from §7a (JSON; `Accept: text/plain` gives a table).
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves, - `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
host health, tokens/s from llama-server's `timings` when present. Scrape it host health, tokens/s from llama-server's `timings` when present. Scrape it
from the fleet Prometheus on orion. from the fleet Prometheus on orion.
## 7a. State store and accounting (SQLite)
One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the
accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64
static build stays a plain `go build`. Single writer goroutine fed by a
channel; readers use their own connection. Volume is a few rows per request,
so nothing here is performance-sensitive.
```sql
CREATE TABLE leases ( -- current table, one row per (route, model)
route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
created INTEGER, last_used INTEGER, PRIMARY KEY (route, model));
CREATE TABLE lease_events ( -- why sessions moved
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
reason TEXT); -- new|unhealthy|idle|pin|release|drain
CREATE TABLE requests ( -- one row per proxied completion
id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT,
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
status INTEGER, streamed INTEGER,
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
err TEXT);
CREATE TABLE host_health ( -- poller observations, for uptime accounting
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
```
Token and cache figures come from the upstream response when llama-server
provides them: `usage` on non-streaming replies, and the final SSE chunk's
`usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on
streamed replies. To see that chunk the proxy tees the response body through
a small SSE line scanner; it never buffers or alters the stream. When the
fields are absent, wall-clock columns are still recorded.
What this answers: per route (session/project), per model, per host — number
of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio
(`cached_tokens / prompt_tokens`, the direct measure of whether affinity is
working), queue wait, error rate, and per-host uptime. `/_crossbar/usage`
exposes the rollups; a nightly job prunes `requests` older than `retention`
(default 180 d) into a `requests_daily` rollup so the file stays small.
Body contents are never stored — only counts and timings.
## 8. Config (`crossbar.yaml`) ## 8. Config (`crossbar.yaml`)
```yaml ```yaml
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0 listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
db: /var/lib/crossbar/crossbar.db
retention: 180d
poll_interval: 60s poll_interval: 60s
lease_idle: 30m lease_idle: 30m
queue_max: 8 queue_max: 8
@@ -157,7 +205,7 @@ Client side, no code changes:
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`. streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
Replaces the hand-maintained provider lists. No leases yet: each route has Replaces the hand-maintained provider lists. No leases yet: each route has
a fixed host list in preference order; first healthy wins. a fixed host list in preference order; first healthy wins.
- **v1:** lease table with persistence, `choose()` by free slots × weight, - **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight,
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ - **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
45 s, else fall through); optional conversation fingerprint (system prompt + 45 s, else fall through); optional conversation fingerprint (system prompt +
@@ -169,6 +217,8 @@ Client side, no code changes:
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming - `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
`/v1/chat/completions` with configurable latency and failure injection. `/v1/chat/completions` with configurable latency and failure injection.
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease. - Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
- SSE tee scanner: fixture streams with and without a final `usage` chunk;
assert the client receives the bytes unchanged and the row is recorded.
- One integration script against the real fleet: one long conversation, - One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server `timings.cache_n` or assert every turn hits the same host (llama-server `timings.cache_n` or
`prompt_n` small after the first turn). `prompt_n` small after the first turn).