PLAN rev 2: SQLite as state store; per-route/model/host accounting
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# crossbar — an affinity router for the fleet's llama-servers
|
||||
|
||||
**Status:** plan, 2026-09-25. Nothing built. Go.
|
||||
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting, Kyle). Nothing built. Go.
|
||||
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
|
||||
|
||||
## 1. Problem
|
||||
@@ -71,8 +71,8 @@ no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
|
||||
lease — the re-prefill is the cost we are avoiding.
|
||||
- **Pinned** leases never move automatically ("project A goes to titan right
|
||||
now"). **Draining** hosts accept no new leases; existing ones finish.
|
||||
- Table persisted to `state.json` after every change; loaded at start so a
|
||||
crossbar restart does not reshuffle sessions.
|
||||
- Table lives in memory and is written through to SQLite (§7a) on every
|
||||
change; loaded at start so a crossbar restart does not reshuffle sessions.
|
||||
|
||||
## 6. Health
|
||||
|
||||
@@ -95,14 +95,62 @@ no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
|
||||
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
|
||||
`{ "release": true }` — drop the lease (next request re-chooses).
|
||||
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
|
||||
- `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups
|
||||
from §7a (JSON; `Accept: text/plain` gives a table).
|
||||
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
|
||||
host health, tokens/s from llama-server's `timings` when present. Scrape it
|
||||
from the fleet Prometheus on orion.
|
||||
|
||||
## 7a. State store and accounting (SQLite)
|
||||
|
||||
One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the
|
||||
accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64
|
||||
static build stays a plain `go build`. Single writer goroutine fed by a
|
||||
channel; readers use their own connection. Volume is a few rows per request,
|
||||
so nothing here is performance-sensitive.
|
||||
|
||||
```sql
|
||||
CREATE TABLE leases ( -- current table, one row per (route, model)
|
||||
route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
|
||||
created INTEGER, last_used INTEGER, PRIMARY KEY (route, model));
|
||||
|
||||
CREATE TABLE lease_events ( -- why sessions moved
|
||||
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
|
||||
reason TEXT); -- new|unhealthy|idle|pin|release|drain
|
||||
|
||||
CREATE TABLE requests ( -- one row per proxied completion
|
||||
id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT,
|
||||
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
|
||||
status INTEGER, streamed INTEGER,
|
||||
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
|
||||
err TEXT);
|
||||
|
||||
CREATE TABLE host_health ( -- poller observations, for uptime accounting
|
||||
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
|
||||
```
|
||||
|
||||
Token and cache figures come from the upstream response when llama-server
|
||||
provides them: `usage` on non-streaming replies, and the final SSE chunk's
|
||||
`usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on
|
||||
streamed replies. To see that chunk the proxy tees the response body through
|
||||
a small SSE line scanner; it never buffers or alters the stream. When the
|
||||
fields are absent, wall-clock columns are still recorded.
|
||||
|
||||
What this answers: per route (session/project), per model, per host — number
|
||||
of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio
|
||||
(`cached_tokens / prompt_tokens`, the direct measure of whether affinity is
|
||||
working), queue wait, error rate, and per-host uptime. `/_crossbar/usage`
|
||||
exposes the rollups; a nightly job prunes `requests` older than `retention`
|
||||
(default 180 d) into a `requests_daily` rollup so the file stays small.
|
||||
|
||||
Body contents are never stored — only counts and timings.
|
||||
|
||||
## 8. Config (`crossbar.yaml`)
|
||||
|
||||
```yaml
|
||||
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
|
||||
db: /var/lib/crossbar/crossbar.db
|
||||
retention: 180d
|
||||
poll_interval: 60s
|
||||
lease_idle: 30m
|
||||
queue_max: 8
|
||||
@@ -157,7 +205,7 @@ Client side, no code changes:
|
||||
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
|
||||
Replaces the hand-maintained provider lists. No leases yet: each route has
|
||||
a fixed host list in preference order; first healthy wins.
|
||||
- **v1:** lease table with persistence, `choose()` by free slots × weight,
|
||||
- **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight,
|
||||
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
|
||||
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
|
||||
45 s, else fall through); optional conversation fingerprint (system prompt +
|
||||
@@ -169,6 +217,8 @@ Client side, no code changes:
|
||||
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
|
||||
`/v1/chat/completions` with configurable latency and failure injection.
|
||||
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
|
||||
- SSE tee scanner: fixture streams with and without a final `usage` chunk;
|
||||
assert the client receives the bytes unchanged and the row is recorded.
|
||||
- One integration script against the real fleet: one long conversation,
|
||||
assert every turn hits the same host (llama-server `timings.cache_n` or
|
||||
`prompt_n` small after the first turn).
|
||||
|
||||
Reference in New Issue
Block a user