diff --git a/PLAN.md b/PLAN.md index 7391daa..4d37973 100644 --- a/PLAN.md +++ b/PLAN.md @@ -1,6 +1,6 @@ # crossbar — an affinity router for the fleet's llama-servers -**Status:** plan, 2026-09-25. Nothing built. Go. +**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting, Kyle). Nothing built. Go. **Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure). ## 1. Problem @@ -71,8 +71,8 @@ no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`, lease — the re-prefill is the cost we are avoiding. - **Pinned** leases never move automatically ("project A goes to titan right now"). **Draining** hosts accept no new leases; existing ones finish. -- Table persisted to `state.json` after every change; loaded at start so a - crossbar restart does not reshuffle sessions. +- Table lives in memory and is written through to SQLite (§7a) on every + change; loaded at start so a crossbar restart does not reshuffle sessions. ## 6. Health @@ -95,14 +95,62 @@ no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`, - `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin; `{ "release": true }` — drop the lease (next request re-chooses). - `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`. +- `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups + from §7a (JSON; `Accept: text/plain` gives a table). - `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves, host health, tokens/s from llama-server's `timings` when present. Scrape it from the fleet Prometheus on orion. +## 7a. State store and accounting (SQLite) + +One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the +accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64 +static build stays a plain `go build`. Single writer goroutine fed by a +channel; readers use their own connection. Volume is a few rows per request, +so nothing here is performance-sensitive. + +```sql +CREATE TABLE leases ( -- current table, one row per (route, model) + route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining + created INTEGER, last_used INTEGER, PRIMARY KEY (route, model)); + +CREATE TABLE lease_events ( -- why sessions moved + ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT, + reason TEXT); -- new|unhealthy|idle|pin|release|drain + +CREATE TABLE requests ( -- one row per proxied completion + id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT, + started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER, + status INTEGER, streamed INTEGER, + prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER, + err TEXT); + +CREATE TABLE host_health ( -- poller observations, for uptime accounting + ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT); +``` + +Token and cache figures come from the upstream response when llama-server +provides them: `usage` on non-streaming replies, and the final SSE chunk's +`usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on +streamed replies. To see that chunk the proxy tees the response body through +a small SSE line scanner; it never buffers or alters the stream. When the +fields are absent, wall-clock columns are still recorded. + +What this answers: per route (session/project), per model, per host — number +of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio +(`cached_tokens / prompt_tokens`, the direct measure of whether affinity is +working), queue wait, error rate, and per-host uptime. `/_crossbar/usage` +exposes the rollups; a nightly job prunes `requests` older than `retention` +(default 180 d) into a `requests_daily` rollup so the file stays small. + +Body contents are never stored — only counts and timings. + ## 8. Config (`crossbar.yaml`) ```yaml listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0 +db: /var/lib/crossbar/crossbar.db +retention: 180d poll_interval: 60s lease_idle: 30m queue_max: 8 @@ -157,7 +205,7 @@ Client side, no code changes: streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins. -- **v1:** lease table with persistence, `choose()` by free slots × weight, +- **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. - **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ 45 s, else fall through); optional conversation fingerprint (system prompt + @@ -169,6 +217,8 @@ Client side, no code changes: - `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming `/v1/chat/completions` with configurable latency and failure injection. - Table tests for `choose()`, lease expiry, drain, passive-health re-lease. +- SSE tee scanner: fixture streams with and without a final `usage` chunk; + assert the client receives the bytes unchanged and the row is recorded. - One integration script against the real fleet: one long conversation, assert every turn hits the same host (llama-server `timings.cache_n` or `prompt_n` small after the first turn).