diff --git a/PLAN.md b/PLAN.md index 4c9d316..f2f3303 100644 --- a/PLAN.md +++ b/PLAN.md @@ -1,6 +1,6 @@ # crossbar — an affinity router for the fleet's llama-servers -**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config — Kyle). Nothing built. Go. +**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go. **Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure). ## 1. Problem @@ -36,8 +36,13 @@ That is the design: **affinity first, balancing only for new sessions.** - **route** — a client identity = the first path segment of the request URL (`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a base URL, so this needs no client support beyond "configurable base URL". -- **lease** — `(route, model) → host`, with `created`, `last_used`, - `state ∈ {active, pinned, draining}`. +- **fingerprint** — a per-conversation key inside a route: SHA-256 of the + request's system prompt plus its first user message (first 4 KiB of each). + Stable across a conversation's turns, different between conversations. Empty + when the body has no user message (probes, title generation). +- **lease** — `(route, fingerprint, model) → host`, with `created`, `last_used`, + `state ∈ {active, pinned, draining}`. Pins are per route; a pinned route + pins all its fingerprints. ## 4. Request flow @@ -45,7 +50,8 @@ That is the design: **affinity first, balancing only for new sessions.** client ── /{route}/v1/chat/completions ──▶ crossbar 1. route := first path segment; strip it 2. model := body.model (JSON peek; fall back to route default) - 3. lease := table[(route, model)] + 2b. fp := fingerprint(body) ("" if none) + 3. lease := table[(route, fp, model)] → fallback table[(route, "", model)] hit & host healthy & model resident → use it miss | host unhealthy → choose(host) ; write lease 4. acquire one concurrency token for (host, model) [bounded queue] @@ -63,6 +69,41 @@ Pass-through per route, all answered from the leased host (or from config if no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`, `/v1/completions`. Everything else 404. +## 4a. Why two keys (Kyle's main use case: many OpenCode instances on one box) + +Per-instance identity comes from the route; per-session identity comes from +the fingerprint. OpenCode's system prompt is per project and its first user +message is per session, so `(route, fp)` separates sessions inside one +instance without any client support. N sessions then spread across hosts by +free slots × weight *at start* and are held there; the per-(host, model) queue +(§6) is what stops N from oversubscribing any one host. + +Client launcher (OpenCode config substitutes `{env:VAR}` in values — verify on +the installed build; if absent, the same value goes through +`options.headers["X-Crossbar-Route"]`, which crossbar also accepts): + +```jsonc +// ~/.config/opencode/opencode.json (one block for every project) +"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible", + "options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" }, + "models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } } +``` +```sh +# oc: one route per instance +CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@" +``` + +## 4b. Context sizes: learned, not configured + +The poller records `n_ctx`, `n_parallel` (→ per-slot context with unified KV) +from `/props` per host and model into `host_health`; `/props` is passed +through per route from the leased host so clients see the real number. Config +may override (`ctx = N` under a model) but normally does not. v2 uses it as a +guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased +host's per-slot context is re-leased to a host where it fits, or answered +400 with a clear message instead of the upstream "Context size has been +exceeded". + ## 5. Lease rules - A lease is **sticky**. It moves only when: the host fails health, the route @@ -111,15 +152,15 @@ so nothing here is performance-sensitive. ```sql CREATE TABLE leases ( -- current table, one row per (route, model) - route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining - created INTEGER, last_used INTEGER, PRIMARY KEY (route, model)); + route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining + created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model)); CREATE TABLE lease_events ( -- why sessions moved ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT, reason TEXT); -- new|unhealthy|idle|pin|release|drain CREATE TABLE requests ( -- one row per proxied completion - id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT, + id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT, started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER, status INTEGER, streamed INTEGER, prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER, @@ -217,18 +258,20 @@ Client side, no code changes: streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins. -- **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight, +- **v1:** SQLite state + accounting (§7a), lease table keyed by + `(route, fingerprint, model)`, header route override, `choose()` by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. - **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ - 45 s, else fall through); optional conversation fingerprint (system prompt + - first user message) for per-conversation stickiness inside one route; - Tailscale-identity route restriction. + 45 s, else fall through); context-size guard (§4b); Tailscale-identity + route restriction. ## 11. Testing - `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming `/v1/chat/completions` with configurable latency and failure injection. -- Table tests for `choose()`, lease expiry, drain, passive-health re-lease. +- Table tests for `choose()`, lease expiry, drain, passive-health re-lease, + fingerprint stability (same conversation → same key; different first user + message → different key; no user message → empty key). - SSE tee scanner: fixture streams with and without a final `usage` chunk; assert the client receives the bytes unchanged and the row is recorded. - One integration script against the real fleet: one long conversation,