PLAN rev 4: per-session leases (route + fingerprint), learned context sizes, many-OpenCode-instances use case

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 00:48:03 -07:00
co-authored by Claude Fable 5.1
parent fe51c20c10
commit 27d43f67ea
+55 -12
View File
@@ -1,6 +1,6 @@
# crossbar — an affinity router for the fleet's llama-servers # crossbar — an affinity router for the fleet's llama-servers
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config — Kyle). Nothing built. Go. **Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go.
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure). **Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
## 1. Problem ## 1. Problem
@@ -36,8 +36,13 @@ That is the design: **affinity first, balancing only for new sessions.**
- **route** — a client identity = the first path segment of the request URL - **route** — a client identity = the first path segment of the request URL
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a (`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
base URL, so this needs no client support beyond "configurable base URL". base URL, so this needs no client support beyond "configurable base URL".
- **lease** — `(route, model) → host`, with `created`, `last_used`, - **fingerprint** — a per-conversation key inside a route: SHA-256 of the
`state ∈ {active, pinned, draining}`. request's system prompt plus its first user message (first 4 KiB of each).
Stable across a conversation's turns, different between conversations. Empty
when the body has no user message (probes, title generation).
- **lease** — `(route, fingerprint, model) → host`, with `created`, `last_used`,
`state ∈ {active, pinned, draining}`. Pins are per route; a pinned route
pins all its fingerprints.
## 4. Request flow ## 4. Request flow
@@ -45,7 +50,8 @@ That is the design: **affinity first, balancing only for new sessions.**
client ── /{route}/v1/chat/completions ──▶ crossbar client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it 1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default) 2. model := body.model (JSON peek; fall back to route default)
3. lease := table[(route, model)] 2b. fp := fingerprint(body) ("" if none)
3. lease := table[(route, fp, model)] → fallback table[(route, "", model)]
hit & host healthy & model resident → use it hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue] 4. acquire one concurrency token for (host, model) [bounded queue]
@@ -63,6 +69,41 @@ Pass-through per route, all answered from the leased host (or from config if
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`, no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
`/v1/completions`. Everything else 404. `/v1/completions`. Everything else 404.
## 4a. Why two keys (Kyle's main use case: many OpenCode instances on one box)
Per-instance identity comes from the route; per-session identity comes from
the fingerprint. OpenCode's system prompt is per project and its first user
message is per session, so `(route, fp)` separates sessions inside one
instance without any client support. N sessions then spread across hosts by
free slots × weight *at start* and are held there; the per-(host, model) queue
(§6) is what stops N from oversubscribing any one host.
Client launcher (OpenCode config substitutes `{env:VAR}` in values — verify on
the installed build; if absent, the same value goes through
`options.headers["X-Crossbar-Route"]`, which crossbar also accepts):
```jsonc
// ~/.config/opencode/opencode.json (one block for every project)
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
"options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" },
"models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } }
```
```sh
# oc: one route per instance
CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@"
```
## 4b. Context sizes: learned, not configured
The poller records `n_ctx`, `n_parallel` (→ per-slot context with unified KV)
from `/props` per host and model into `host_health`; `/props` is passed
through per route from the leased host so clients see the real number. Config
may override (`ctx = N` under a model) but normally does not. v2 uses it as a
guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased
host's per-slot context is re-leased to a host where it fits, or answered
400 with a clear message instead of the upstream "Context size has been
exceeded".
## 5. Lease rules ## 5. Lease rules
- A lease is **sticky**. It moves only when: the host fails health, the route - A lease is **sticky**. It moves only when: the host fails health, the route
@@ -111,15 +152,15 @@ so nothing here is performance-sensitive.
```sql ```sql
CREATE TABLE leases ( -- current table, one row per (route, model) CREATE TABLE leases ( -- current table, one row per (route, model)
route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
created INTEGER, last_used INTEGER, PRIMARY KEY (route, model)); created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model));
CREATE TABLE lease_events ( -- why sessions moved CREATE TABLE lease_events ( -- why sessions moved
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT, ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
reason TEXT); -- new|unhealthy|idle|pin|release|drain reason TEXT); -- new|unhealthy|idle|pin|release|drain
CREATE TABLE requests ( -- one row per proxied completion CREATE TABLE requests ( -- one row per proxied completion
id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT, id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT,
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER, started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
status INTEGER, streamed INTEGER, status INTEGER, streamed INTEGER,
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER, prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
@@ -217,18 +258,20 @@ Client side, no code changes:
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`. streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
Replaces the hand-maintained provider lists. No leases yet: each route has Replaces the hand-maintained provider lists. No leases yet: each route has
a fixed host list in preference order; first healthy wins. a fixed host list in preference order; first healthy wins.
- **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight, - **v1:** SQLite state + accounting (§7a), lease table keyed by
`(route, fingerprint, model)`, header route override, `choose()` by free slots × weight,
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ - **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
45 s, else fall through); optional conversation fingerprint (system prompt + 45 s, else fall through); context-size guard (§4b); Tailscale-identity
first user message) for per-conversation stickiness inside one route; route restriction.
Tailscale-identity route restriction.
## 11. Testing ## 11. Testing
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming - `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
`/v1/chat/completions` with configurable latency and failure injection. `/v1/chat/completions` with configurable latency and failure injection.
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease. - Table tests for `choose()`, lease expiry, drain, passive-health re-lease,
fingerprint stability (same conversation → same key; different first user
message → different key; no user message → empty key).
- SSE tee scanner: fixture streams with and without a final `usage` chunk; - SSE tee scanner: fixture streams with and without a final `usage` chunk;
assert the client receives the bytes unchanged and the row is recorded. assert the client receives the bytes unchanged and the row is recorded.
- One integration script against the real fleet: one long conversation, - One integration script against the real fleet: one long conversation,