12 KiB
crossbar — an affinity router for the fleet's llama-servers
Status: plan, 2026-09-25 (rev 2: SQLite state + accounting, Kyle). Nothing built. Go. Owner: Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
1. Problem
Three hosts run llama-server in router mode with the same model ids
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k tokens with ~100 % cache hits. Any move between hosts re-prefills the whole context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV ("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and holds the connection. That is the design: affinity first, balancing only for new sessions.
2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts. The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (
--models-max). crossbar never sends a request that would make a host load a model it does not have resident. - No auth in v0 (tailnet-only bind). See §9.
3. Terminology
- host — one llama-server router endpoint (
base_url, list of model ids it can serve, per-modelparallel, a speedweight). - route — a client identity = the first path segment of the request URL
(
/opencode-a/v1/...→ routeopencode-a). The client only ever knows a base URL, so this needs no client support beyond "configurable base URL". - lease —
(route, model) → host, withcreated,last_used,state ∈ {active, pinned, draining}.
4. Request flow
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
3. lease := table[(route, model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
choose(host): among hosts that are healthy and list the model as loaded
(/models), take the one with the most free slots × weight; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that can
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): /v1/models, /health, /props, /v1/embeddings,
/v1/completions. Everything else 404.
5. Lease rules
- A lease is sticky. It moves only when: the host fails health, the route
is idle longer than
lease_idle(default 30 min), or an operator pins/ releases it. A faster host coming back online does not move an active lease — the re-prefill is the cost we are avoiding. - Pinned leases never move automatically ("project A goes to titan right now"). Draining hosts accept no new leases; existing ones finish.
- Table lives in memory and is written through to SQLite (§7a) on every change; loaded at start so a crossbar restart does not reshuffle sessions.
6. Health
- Poller, every
poll_interval(60 s):GET /healththenGET /modelsper host. Record loaded models. Only thenGET /slots?model=Xfor loaded models only — probing an unloaded model makes llama-server load it. Per-model free-slot count from/slots(or, if/slotsis disabled on a host, assumeparallelminus our own in-flight count). - Passive: a connection error or 5xx on a proxied request marks the host unhealthy immediately and re-leases the route on the client's retry. Recovery only through the poller (two consecutive good polls).
- Concurrency: per
(host, model)token bucket of sizeparallel. Requests beyond it wait in a bounded FIFO (queue_max, default 8; 503 beyond that). This is where the unified-KV overflow is prevented.
7. Admin API (tailnet only, same listener, prefix /_crossbar)
GET /_crossbar/hosts— health, loaded models, free slots, in-flight.GET /_crossbar/routes— the lease table.POST /_crossbar/routes/{route}{ "host": "titan", "pin": true }— pin;{ "release": true }— drop the lease (next request re-chooses).POST /_crossbar/hosts/{host}{ "drain": true|false }.GET /_crossbar/usage?since=…&by=route|model|host— accounting rollups from §7a (JSON;Accept: text/plaingives a table).GET /_crossbar/metrics— Prometheus: requests, queue wait, lease moves, host health, tokens/s from llama-server'stimingswhen present. Scrape it from the fleet Prometheus on orion.
7a. State store and accounting (SQLite)
One SQLite file (crossbar.db, WAL mode) holds both the durable state and the
accounting log. Driver: modernc.org/sqlite (pure Go, no cgo) so the arm64
static build stays a plain go build. Single writer goroutine fed by a
channel; readers use their own connection. Volume is a few rows per request,
so nothing here is performance-sensitive.
CREATE TABLE leases ( -- current table, one row per (route, model)
route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
created INTEGER, last_used INTEGER, PRIMARY KEY (route, model));
CREATE TABLE lease_events ( -- why sessions moved
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
reason TEXT); -- new|unhealthy|idle|pin|release|drain
CREATE TABLE requests ( -- one row per proxied completion
id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT,
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
status INTEGER, streamed INTEGER,
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
err TEXT);
CREATE TABLE host_health ( -- poller observations, for uptime accounting
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
Token and cache figures come from the upstream response when llama-server
provides them: usage on non-streaming replies, and the final SSE chunk's
usage / timings (prompt_n, cache_n, predicted_n, predicted_ms) on
streamed replies. To see that chunk the proxy tees the response body through
a small SSE line scanner; it never buffers or alters the stream. When the
fields are absent, wall-clock columns are still recorded.
What this answers: per route (session/project), per model, per host — number
of requests, busy seconds (sum(total_ms)), tokens in/out, cache-hit ratio
(cached_tokens / prompt_tokens, the direct measure of whether affinity is
working), queue wait, error rate, and per-host uptime. /_crossbar/usage
exposes the rollups; a nightly job prunes requests older than retention
(default 180 d) into a requests_daily rollup so the file stays small.
Body contents are never stored — only counts and timings.
8. Config (crossbar.yaml)
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
db: /var/lib/crossbar/crossbar.db
retention: 180d
poll_interval: 60s
lease_idle: 30m
queue_max: 8
hosts:
straylight:
base_url: http://straylight.<tailnet>:11434
weight: 1.0
models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
titan:
base_url: http://titan.<tailnet>:8081
weight: 2.0 # ~2× straylight decode
models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
dixie:
base_url: http://dixie.<tailnet>:11434
weight: 0.8
models: { ornith-1.5-9b-uncensored: {parallel: 6} }
routes: # optional defaults / pins
opencode-a: { default_model: ornith-1.5-35b-a3b }
opencode-b: { default_model: ornith-1.5-35b-a3b }
paper: { default_model: qwen3.8-27b-uncensored, pin: titan }
hermes-straylight: {}
hermes-talos: {}
hermes-titan: { pin: titan } # titan's models are the titan agent's first
Client side, no code changes:
- OpenCode: project-local
opencode.jsonprovider withoptions.baseURL: http://crossbar:7777/opencode-a/v1(global + project configs merge, so each project carries its own route). - Hermes:
custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1(and thedelegation/auxiliaryblocks that point at a router today). - tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (
plain_http_to_sinkfor that host only).
9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (
tailscale whoison the peer address) and restrict routes to peers, sohermes-taloscan only be used from talos. - Admin API: same listener, same trust. Consider a separate
admin_listenon localhost if crossbar runs on a shared host. - No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and timings only.
10. Milestones
- v0 (a day): config, static routes → host mapping, reverse proxy with
streaming,
/health+/modelspoller, passive health,/_crossbar/hosts. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins. - v1: SQLite state + accounting (§7a), lease table,
choose()by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. - v2: wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ 45 s, else fall through); optional conversation fingerprint (system prompt + first user message) for per-conversation stickiness inside one route; Tailscale-identity route restriction.
11. Testing
httptestfake llama-server:/health,/models,/slots, streaming/v1/chat/completionswith configurable latency and failure injection.- Table tests for
choose(), lease expiry, drain, passive-health re-lease. - SSE tee scanner: fixture streams with and without a final
usagechunk; assert the client receives the bytes unchanged and the row is recorded. - One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server
timings.cache_norprompt_nsmall after the first turn).
12. Open questions
- Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static Go binary + systemd unit; arm64 cross-compile is free. Candidate: hyperborea, in its own unit, outside the collector's disk budget.
/slotsis disabled by default on recent llama-server builds; enable--slotson the three routers or rely on our own in-flight counts.- Should
default_modelrewrite a client'smodelfield? Proposal: no — proxy what the client sent; only use the default when the body has none. - Does the titan agent's reservation (2026-09-21) stand as a pin, or does titan become a general pool member now that health is tracked? Kyle said titan could join the pool (2026-09-25); the pin above keeps its own agent first without excluding others.