PLAN rev 4: per-session leases (route + fingerprint), learned context sizes, many-OpenCode-instances use case
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# crossbar — an affinity router for the fleet's llama-servers
|
||||
|
||||
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config — Kyle). Nothing built. Go.
|
||||
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go.
|
||||
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
|
||||
|
||||
## 1. Problem
|
||||
@@ -36,8 +36,13 @@ That is the design: **affinity first, balancing only for new sessions.**
|
||||
- **route** — a client identity = the first path segment of the request URL
|
||||
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
|
||||
base URL, so this needs no client support beyond "configurable base URL".
|
||||
- **lease** — `(route, model) → host`, with `created`, `last_used`,
|
||||
`state ∈ {active, pinned, draining}`.
|
||||
- **fingerprint** — a per-conversation key inside a route: SHA-256 of the
|
||||
request's system prompt plus its first user message (first 4 KiB of each).
|
||||
Stable across a conversation's turns, different between conversations. Empty
|
||||
when the body has no user message (probes, title generation).
|
||||
- **lease** — `(route, fingerprint, model) → host`, with `created`, `last_used`,
|
||||
`state ∈ {active, pinned, draining}`. Pins are per route; a pinned route
|
||||
pins all its fingerprints.
|
||||
|
||||
## 4. Request flow
|
||||
|
||||
@@ -45,7 +50,8 @@ That is the design: **affinity first, balancing only for new sessions.**
|
||||
client ── /{route}/v1/chat/completions ──▶ crossbar
|
||||
1. route := first path segment; strip it
|
||||
2. model := body.model (JSON peek; fall back to route default)
|
||||
3. lease := table[(route, model)]
|
||||
2b. fp := fingerprint(body) ("" if none)
|
||||
3. lease := table[(route, fp, model)] → fallback table[(route, "", model)]
|
||||
hit & host healthy & model resident → use it
|
||||
miss | host unhealthy → choose(host) ; write lease
|
||||
4. acquire one concurrency token for (host, model) [bounded queue]
|
||||
@@ -63,6 +69,41 @@ Pass-through per route, all answered from the leased host (or from config if
|
||||
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
|
||||
`/v1/completions`. Everything else 404.
|
||||
|
||||
## 4a. Why two keys (Kyle's main use case: many OpenCode instances on one box)
|
||||
|
||||
Per-instance identity comes from the route; per-session identity comes from
|
||||
the fingerprint. OpenCode's system prompt is per project and its first user
|
||||
message is per session, so `(route, fp)` separates sessions inside one
|
||||
instance without any client support. N sessions then spread across hosts by
|
||||
free slots × weight *at start* and are held there; the per-(host, model) queue
|
||||
(§6) is what stops N from oversubscribing any one host.
|
||||
|
||||
Client launcher (OpenCode config substitutes `{env:VAR}` in values — verify on
|
||||
the installed build; if absent, the same value goes through
|
||||
`options.headers["X-Crossbar-Route"]`, which crossbar also accepts):
|
||||
|
||||
```jsonc
|
||||
// ~/.config/opencode/opencode.json (one block for every project)
|
||||
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
|
||||
"options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" },
|
||||
"models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } }
|
||||
```
|
||||
```sh
|
||||
# oc: one route per instance
|
||||
CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@"
|
||||
```
|
||||
|
||||
## 4b. Context sizes: learned, not configured
|
||||
|
||||
The poller records `n_ctx`, `n_parallel` (→ per-slot context with unified KV)
|
||||
from `/props` per host and model into `host_health`; `/props` is passed
|
||||
through per route from the leased host so clients see the real number. Config
|
||||
may override (`ctx = N` under a model) but normally does not. v2 uses it as a
|
||||
guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased
|
||||
host's per-slot context is re-leased to a host where it fits, or answered
|
||||
400 with a clear message instead of the upstream "Context size has been
|
||||
exceeded".
|
||||
|
||||
## 5. Lease rules
|
||||
|
||||
- A lease is **sticky**. It moves only when: the host fails health, the route
|
||||
@@ -111,15 +152,15 @@ so nothing here is performance-sensitive.
|
||||
|
||||
```sql
|
||||
CREATE TABLE leases ( -- current table, one row per (route, model)
|
||||
route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
|
||||
created INTEGER, last_used INTEGER, PRIMARY KEY (route, model));
|
||||
route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
|
||||
created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model));
|
||||
|
||||
CREATE TABLE lease_events ( -- why sessions moved
|
||||
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
|
||||
reason TEXT); -- new|unhealthy|idle|pin|release|drain
|
||||
|
||||
CREATE TABLE requests ( -- one row per proxied completion
|
||||
id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT,
|
||||
id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT,
|
||||
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
|
||||
status INTEGER, streamed INTEGER,
|
||||
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
|
||||
@@ -217,18 +258,20 @@ Client side, no code changes:
|
||||
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
|
||||
Replaces the hand-maintained provider lists. No leases yet: each route has
|
||||
a fixed host list in preference order; first healthy wins.
|
||||
- **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight,
|
||||
- **v1:** SQLite state + accounting (§7a), lease table keyed by
|
||||
`(route, fingerprint, model)`, header route override, `choose()` by free slots × weight,
|
||||
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
|
||||
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
|
||||
45 s, else fall through); optional conversation fingerprint (system prompt +
|
||||
first user message) for per-conversation stickiness inside one route;
|
||||
Tailscale-identity route restriction.
|
||||
45 s, else fall through); context-size guard (§4b); Tailscale-identity
|
||||
route restriction.
|
||||
|
||||
## 11. Testing
|
||||
|
||||
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
|
||||
`/v1/chat/completions` with configurable latency and failure injection.
|
||||
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
|
||||
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease,
|
||||
fingerprint stability (same conversation → same key; different first user
|
||||
message → different key; no user message → empty key).
|
||||
- SSE tee scanner: fixture streams with and without a final `usage` chunk;
|
||||
assert the client receives the bytes unchanged and the row is recorded.
|
||||
- One integration script against the real fleet: one long conversation,
|
||||
|
||||
Reference in New Issue
Block a user