# crossbar — an affinity router for the fleet's llama-servers **Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go. **Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure). ## 1. Problem Three hosts run `llama-server` in router mode with the *same model ids* (straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents, the paper writer) each hard-code one host. Consequences seen this week: - a host's prompt cache is per process; an agent session sits at 100–230k tokens with ~100 % cache hits. Any move between hosts re-prefills the whole context (≈ 7–8 min at straylight's ~500 t/s prompt processing); - two turns landing on one 35B at once overflowed its unified KV ("Context size has been exceeded"); nothing in the path knows the slot count; - titan sleeps; clients pointed at it fail instead of falling back; - rebuilding or restarting a host takes its clients down with it. A crossbar switch connects each caller to a line and *holds the connection*. That is the design: **affinity first, balancing only for new sessions.** ## 2. Non-goals - No token counting, prompt rewriting, batching or splitting across hosts. The hosts are bandwidth-bound; there is nothing to gain cross-host. - No model management. Loading/unloading stays with each host's llama-server router (`--models-max`). crossbar never sends a request that would make a host load a model it does not have resident. - No auth in v0 (tailnet-only bind). See §9. ## 3. Terminology - **host** — one llama-server router endpoint (`base_url`, list of model ids it can serve, per-model `parallel`, a speed `weight`). - **route** — a client identity = the first path segment of the request URL (`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a base URL, so this needs no client support beyond "configurable base URL". - **fingerprint** — a per-conversation key inside a route: SHA-256 of the request's system prompt plus its first user message (first 4 KiB of each). Stable across a conversation's turns, different between conversations. Empty when the body has no user message (probes, title generation). - **lease** — `(route, fingerprint, model) → host`, with `created`, `last_used`, `state ∈ {active, pinned, draining}`. Pins are per route; a pinned route pins all its fingerprints. ## 4. Request flow ``` client ── /{route}/v1/chat/completions ──▶ crossbar 1. route := first path segment; strip it 2. model := body.model (JSON peek; fall back to route default) 2b. fp := fingerprint(body) ("" if none) 3. lease := table[(route, fp, model)] → fallback table[(route, "", model)] hit & host healthy & model resident → use it miss | host unhealthy → choose(host) ; write lease 4. acquire one concurrency token for (host, model) [bounded queue] 5. reverse-proxy to host.base_url + "/v1/..." ; stream through, flush on every chunk; passive health on connect error / 5xx 6. release token; lease.last_used = now ``` `choose(host)`: among hosts that are healthy **and list the model as loaded** (`/models`), take the one with the most free slots × `weight`; tie → lowest current queue depth. If no host has it loaded, take a healthy host that *can* serve it (config) and accept the load; log that decision. Pass-through per route, all answered from the leased host (or from config if no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`, `/v1/completions`. Everything else 404. ## 4a. Why two keys (Kyle's main use case: many OpenCode instances on one box) Per-instance identity comes from the route; per-session identity comes from the fingerprint. OpenCode's system prompt is per project and its first user message is per session, so `(route, fp)` separates sessions inside one instance without any client support. N sessions then spread across hosts by free slots × weight *at start* and are held there; the per-(host, model) queue (§6) is what stops N from oversubscribing any one host. Client launcher (OpenCode config substitutes `{env:VAR}` in values — verified in the installed 1.15.10 bundle: `/\{env:([^}]+)\}/g` → `process.env[V]`, missing → empty string, so always set the variable; as a fallback the same value can go through `options.headers["X-Crossbar-Route"]`, which crossbar also accepts): ```jsonc // ~/.config/opencode/opencode.json (one block for every project) "provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible", "options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" }, "models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } } ``` ```sh # oc: one route per instance CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@" ``` ## 4b. Context sizes: learned, not configured The poller records `n_ctx`, `n_parallel` (→ per-slot context with unified KV) from `/props` per host and model into `host_health`; `/props` is passed through per route from the leased host so clients see the real number. Config may override (`ctx = N` under a model) but normally does not. v2 uses it as a guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased host's per-slot context is re-leased to a host where it fits, or answered 400 with a clear message instead of the upstream "Context size has been exceeded". ## 5. Lease rules - A lease is **sticky**. It moves only when: the host fails health, the route is idle longer than `lease_idle` (default 30 min), or an operator pins/ releases it. A faster host coming back online does *not* move an active lease — the re-prefill is the cost we are avoiding. - **Pinned** leases never move automatically ("project A goes to titan right now"). **Draining** hosts accept no new leases; existing ones finish. - Table lives in memory and is written through to SQLite (§7a) on every change; loaded at start so a crossbar restart does not reshuffle sessions. ## 6. Health - Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per host. Record loaded models. Only then `GET /slots?model=X` **for loaded models only** — probing an unloaded model makes llama-server load it. Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a host, assume `parallel` minus our own in-flight count). - Passive: a connection error or 5xx on a proxied request marks the host unhealthy immediately and re-leases the route on the client's retry. Recovery only through the poller (two consecutive good polls). - Concurrency: per `(host, model)` token bucket of size `parallel`. Requests beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that). This is where the unified-KV overflow is prevented. ## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`) - `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight. - `GET /_crossbar/routes` — the lease table. - `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin; `{ "release": true }` — drop the lease (next request re-chooses). - `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`. - `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups from §7a (JSON; `Accept: text/plain` gives a table). - `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves, host health, tokens/s from llama-server's `timings` when present. Scrape it from the fleet Prometheus on orion. ## 7a. State store and accounting (SQLite) One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64 static build stays a plain `go build`. Single writer goroutine fed by a channel; readers use their own connection. Volume is a few rows per request, so nothing here is performance-sensitive. ```sql CREATE TABLE leases ( -- current table, one row per (route, model) route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model)); CREATE TABLE lease_events ( -- why sessions moved ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT, reason TEXT); -- new|unhealthy|idle|pin|release|drain CREATE TABLE requests ( -- one row per proxied completion id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT, started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER, status INTEGER, streamed INTEGER, prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER, err TEXT); CREATE TABLE host_health ( -- poller observations, for uptime accounting ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT); ``` Token and cache figures come from the upstream response when llama-server provides them: `usage` on non-streaming replies, and the final SSE chunk's `usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on streamed replies. To see that chunk the proxy tees the response body through a small SSE line scanner; it never buffers or alters the stream. When the fields are absent, wall-clock columns are still recorded. What this answers: per route (session/project), per model, per host — number of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio (`cached_tokens / prompt_tokens`, the direct measure of whether affinity is working), queue wait, error rate, and per-host uptime. `/_crossbar/usage` exposes the rollups; a nightly job prunes `requests` older than `retention` (default 180 d) into a `requests_daily` rollup so the file stays small. Body contents are never stored — only counts and timings. ## 8. Config (`crossbar.toml`) TOML (`github.com/BurntSushi/toml`): no indentation semantics, no implicit type coercion, and it is what Kyle's Rust projects already use. A NixOS module can generate it with `pkgs.formats.toml`. ```toml listen = "100.x.y.z:7777" # tailnet address only; never 0.0.0.0 db = "/var/lib/crossbar/crossbar.db" poll_interval = "60s" lease_idle = "30m" queue_max = 8 retention = "180d" [hosts.straylight] base_url = "http://straylight.:11434" weight = 1.0 models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "ornith-1.5-9b-uncensored" = { parallel = 6 } } [hosts.titan] base_url = "http://titan.:8081" weight = 2.0 # ~2x straylight decode models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "laguna-s-2.1" = { parallel = 2 } } [hosts.dixie] base_url = "http://dixie.:11434" weight = 0.8 models = { "ornith-1.5-9b-uncensored" = { parallel = 6 } } # optional per-route defaults / pins [routes.opencode-a] default_model = "ornith-1.5-35b-a3b" [routes.opencode-b] default_model = "ornith-1.5-35b-a3b" [routes.paper] default_model = "qwen3.8-27b-uncensored" pin = "titan" [routes.hermes-straylight] [routes.hermes-talos] [routes.hermes-titan] pin = "titan" # titan's models are the titan agent's first ``` Client side, no code changes: - OpenCode: project-local `opencode.json` provider with `options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project configs merge, so each project carries its own route). - Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-/v1` (and the `delegation` / `auxiliary` blocks that point at a router today). - tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry the llama routers already have (`plain_http_to_sink` for that host only). ## 9. Security notes - Bind to the tailnet address only. v0 relies on the tailnet for authentication; every route is reachable by every tailnet peer, which is the same exposure the llama-servers have today. v1 option: read Tailscale identity (`tailscale whois` on the peer address) and restrict routes to peers, so `hermes-talos` can only be used from talos. - Admin API: same listener, same trust. Consider a separate `admin_listen` on localhost if crossbar runs on a shared host. - No secrets in config; llama-servers take no keys. - Request bodies are proxied, never logged. Metrics carry counts and timings only. ## 10. Milestones - **v0 (a day):** TOML config, static routes → host mapping, reverse proxy with streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins. - **v1:** SQLite state + accounting (§7a), lease table keyed by `(route, fingerprint, model)`, header route override, `choose()` by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. - **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ 45 s, else fall through); context-size guard (§4b); Tailscale-identity route restriction. ## 11. Testing - `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming `/v1/chat/completions` with configurable latency and failure injection. - Table tests for `choose()`, lease expiry, drain, passive-health re-lease, fingerprint stability (same conversation → same key; different first user message → different key; no user message → empty key). - SSE tee scanner: fixture streams with and without a final `usage` chunk; assert the client receives the bytes unchanged and the row is recorded. - One integration script against the real fleet: one long conversation, assert every turn hits the same host (llama-server `timings.cache_n` or `prompt_n` small after the first turn). ## 12. Open questions 1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static Go binary + systemd unit; arm64 cross-compile is free. Candidate: hyperborea, in its own unit, outside the collector's disk budget. 2. `/slots` is disabled by default on recent llama-server builds; enable `--slots` on the three routers or rely on our own in-flight counts. 3. Should `default_model` rewrite a client's `model` field? Proposal: no — proxy what the client sent; only use the default when the body has none. 4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does titan become a general pool member now that health is tracked? Kyle said titan could join the pool (2026-09-25); the pin above keeps its own agent first without excluding others.