14 KiB
crossbar — an affinity router for the fleet's llama-servers
Status: plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go. Owner: Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
1. Problem
Three hosts run llama-server in router mode with the same model ids
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k tokens with ~100 % cache hits. Any move between hosts re-prefills the whole context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV ("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and holds the connection. That is the design: affinity first, balancing only for new sessions.
2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts. The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (
--models-max). crossbar never sends a request that would make a host load a model it does not have resident. - No auth in v0 (tailnet-only bind). See §9.
3. Terminology
- host — one llama-server router endpoint (
base_url, list of model ids it can serve, per-modelparallel, a speedweight). - route — a client identity = the first path segment of the request URL
(
/opencode-a/v1/...→ routeopencode-a). The client only ever knows a base URL, so this needs no client support beyond "configurable base URL". - fingerprint — a per-conversation key inside a route: SHA-256 of the request's system prompt plus its first user message (first 4 KiB of each). Stable across a conversation's turns, different between conversations. Empty when the body has no user message (probes, title generation).
- lease —
(route, fingerprint, model) → host, withcreated,last_used,state ∈ {active, pinned, draining}. Pins are per route; a pinned route pins all its fingerprints.
4. Request flow
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
2b. fp := fingerprint(body) ("" if none)
3. lease := table[(route, fp, model)] → fallback table[(route, "", model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
choose(host): among hosts that are healthy and list the model as loaded
(/models), take the one with the most free slots × weight; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that can
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): /v1/models, /health, /props, /v1/embeddings,
/v1/completions. Everything else 404.
4a. Why two keys (Kyle's main use case: many OpenCode instances on one box)
Per-instance identity comes from the route; per-session identity comes from
the fingerprint. OpenCode's system prompt is per project and its first user
message is per session, so (route, fp) separates sessions inside one
instance without any client support. N sessions then spread across hosts by
free slots × weight at start and are held there; the per-(host, model) queue
(§6) is what stops N from oversubscribing any one host.
Client launcher (OpenCode config substitutes {env:VAR} in values — verify on
the installed build; if absent, the same value goes through
options.headers["X-Crossbar-Route"], which crossbar also accepts):
// ~/.config/opencode/opencode.json (one block for every project)
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
"options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" },
"models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } }
# oc: one route per instance
CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@"
4b. Context sizes: learned, not configured
The poller records n_ctx, n_parallel (→ per-slot context with unified KV)
from /props per host and model into host_health; /props is passed
through per route from the leased host so clients see the real number. Config
may override (ctx = N under a model) but normally does not. v2 uses it as a
guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased
host's per-slot context is re-leased to a host where it fits, or answered
400 with a clear message instead of the upstream "Context size has been
exceeded".
5. Lease rules
- A lease is sticky. It moves only when: the host fails health, the route
is idle longer than
lease_idle(default 30 min), or an operator pins/ releases it. A faster host coming back online does not move an active lease — the re-prefill is the cost we are avoiding. - Pinned leases never move automatically ("project A goes to titan right now"). Draining hosts accept no new leases; existing ones finish.
- Table lives in memory and is written through to SQLite (§7a) on every change; loaded at start so a crossbar restart does not reshuffle sessions.
6. Health
- Poller, every
poll_interval(60 s):GET /healththenGET /modelsper host. Record loaded models. Only thenGET /slots?model=Xfor loaded models only — probing an unloaded model makes llama-server load it. Per-model free-slot count from/slots(or, if/slotsis disabled on a host, assumeparallelminus our own in-flight count). - Passive: a connection error or 5xx on a proxied request marks the host unhealthy immediately and re-leases the route on the client's retry. Recovery only through the poller (two consecutive good polls).
- Concurrency: per
(host, model)token bucket of sizeparallel. Requests beyond it wait in a bounded FIFO (queue_max, default 8; 503 beyond that). This is where the unified-KV overflow is prevented.
7. Admin API (tailnet only, same listener, prefix /_crossbar)
GET /_crossbar/hosts— health, loaded models, free slots, in-flight.GET /_crossbar/routes— the lease table.POST /_crossbar/routes/{route}{ "host": "titan", "pin": true }— pin;{ "release": true }— drop the lease (next request re-chooses).POST /_crossbar/hosts/{host}{ "drain": true|false }.GET /_crossbar/usage?since=…&by=route|model|host— accounting rollups from §7a (JSON;Accept: text/plaingives a table).GET /_crossbar/metrics— Prometheus: requests, queue wait, lease moves, host health, tokens/s from llama-server'stimingswhen present. Scrape it from the fleet Prometheus on orion.
7a. State store and accounting (SQLite)
One SQLite file (crossbar.db, WAL mode) holds both the durable state and the
accounting log. Driver: modernc.org/sqlite (pure Go, no cgo) so the arm64
static build stays a plain go build. Single writer goroutine fed by a
channel; readers use their own connection. Volume is a few rows per request,
so nothing here is performance-sensitive.
CREATE TABLE leases ( -- current table, one row per (route, model)
route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model));
CREATE TABLE lease_events ( -- why sessions moved
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
reason TEXT); -- new|unhealthy|idle|pin|release|drain
CREATE TABLE requests ( -- one row per proxied completion
id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT,
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
status INTEGER, streamed INTEGER,
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
err TEXT);
CREATE TABLE host_health ( -- poller observations, for uptime accounting
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
Token and cache figures come from the upstream response when llama-server
provides them: usage on non-streaming replies, and the final SSE chunk's
usage / timings (prompt_n, cache_n, predicted_n, predicted_ms) on
streamed replies. To see that chunk the proxy tees the response body through
a small SSE line scanner; it never buffers or alters the stream. When the
fields are absent, wall-clock columns are still recorded.
What this answers: per route (session/project), per model, per host — number
of requests, busy seconds (sum(total_ms)), tokens in/out, cache-hit ratio
(cached_tokens / prompt_tokens, the direct measure of whether affinity is
working), queue wait, error rate, and per-host uptime. /_crossbar/usage
exposes the rollups; a nightly job prunes requests older than retention
(default 180 d) into a requests_daily rollup so the file stays small.
Body contents are never stored — only counts and timings.
8. Config (crossbar.toml)
TOML (github.com/BurntSushi/toml): no indentation semantics, no implicit
type coercion, and it is what Kyle's Rust projects already use. A NixOS module
can generate it with pkgs.formats.toml.
listen = "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
db = "/var/lib/crossbar/crossbar.db"
poll_interval = "60s"
lease_idle = "30m"
queue_max = 8
retention = "180d"
[hosts.straylight]
base_url = "http://straylight.<tailnet>:11434"
weight = 1.0
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "ornith-1.5-9b-uncensored" = { parallel = 6 } }
[hosts.titan]
base_url = "http://titan.<tailnet>:8081"
weight = 2.0 # ~2x straylight decode
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "laguna-s-2.1" = { parallel = 2 } }
[hosts.dixie]
base_url = "http://dixie.<tailnet>:11434"
weight = 0.8
models = { "ornith-1.5-9b-uncensored" = { parallel = 6 } }
# optional per-route defaults / pins
[routes.opencode-a]
default_model = "ornith-1.5-35b-a3b"
[routes.opencode-b]
default_model = "ornith-1.5-35b-a3b"
[routes.paper]
default_model = "qwen3.8-27b-uncensored"
pin = "titan"
[routes.hermes-straylight]
[routes.hermes-talos]
[routes.hermes-titan]
pin = "titan" # titan's models are the titan agent's first
Client side, no code changes:
- OpenCode: project-local
opencode.jsonprovider withoptions.baseURL: http://crossbar:7777/opencode-a/v1(global + project configs merge, so each project carries its own route). - Hermes:
custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1(and thedelegation/auxiliaryblocks that point at a router today). - tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (
plain_http_to_sinkfor that host only).
9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (
tailscale whoison the peer address) and restrict routes to peers, sohermes-taloscan only be used from talos. - Admin API: same listener, same trust. Consider a separate
admin_listenon localhost if crossbar runs on a shared host. - No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and timings only.
10. Milestones
- v0 (a day): TOML config, static routes → host mapping, reverse proxy with
streaming,
/health+/modelspoller, passive health,/_crossbar/hosts. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins. - v1: SQLite state + accounting (§7a), lease table keyed by
(route, fingerprint, model), header route override,choose()by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. - v2: wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ 45 s, else fall through); context-size guard (§4b); Tailscale-identity route restriction.
11. Testing
httptestfake llama-server:/health,/models,/slots, streaming/v1/chat/completionswith configurable latency and failure injection.- Table tests for
choose(), lease expiry, drain, passive-health re-lease, fingerprint stability (same conversation → same key; different first user message → different key; no user message → empty key). - SSE tee scanner: fixture streams with and without a final
usagechunk; assert the client receives the bytes unchanged and the row is recorded. - One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server
timings.cache_norprompt_nsmall after the first turn).
12. Open questions
- Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static Go binary + systemd unit; arm64 cross-compile is free. Candidate: hyperborea, in its own unit, outside the collector's disk budget.
/slotsis disabled by default on recent llama-server builds; enable--slotson the three routers or rely on our own in-flight counts.- Should
default_modelrewrite a client'smodelfield? Proposal: no — proxy what the client sent; only use the default when the body has none. - Does the titan agent's reservation (2026-09-21) stand as a pin, or does titan become a general pool member now that health is tracked? Kyle said titan could join the pool (2026-09-25); the pin above keeps its own agent first without excluding others.