9.0 KiB
crossbar — an affinity router for the fleet's llama-servers
Status: plan, 2026-09-25. Nothing built. Go. Owner: Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
1. Problem
Three hosts run llama-server in router mode with the same model ids
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k tokens with ~100 % cache hits. Any move between hosts re-prefills the whole context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV ("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and holds the connection. That is the design: affinity first, balancing only for new sessions.
2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts. The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (
--models-max). crossbar never sends a request that would make a host load a model it does not have resident. - No auth in v0 (tailnet-only bind). See §9.
3. Terminology
- host — one llama-server router endpoint (
base_url, list of model ids it can serve, per-modelparallel, a speedweight). - route — a client identity = the first path segment of the request URL
(
/opencode-a/v1/...→ routeopencode-a). The client only ever knows a base URL, so this needs no client support beyond "configurable base URL". - lease —
(route, model) → host, withcreated,last_used,state ∈ {active, pinned, draining}.
4. Request flow
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
3. lease := table[(route, model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
choose(host): among hosts that are healthy and list the model as loaded
(/models), take the one with the most free slots × weight; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that can
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): /v1/models, /health, /props, /v1/embeddings,
/v1/completions. Everything else 404.
5. Lease rules
- A lease is sticky. It moves only when: the host fails health, the route
is idle longer than
lease_idle(default 30 min), or an operator pins/ releases it. A faster host coming back online does not move an active lease — the re-prefill is the cost we are avoiding. - Pinned leases never move automatically ("project A goes to titan right now"). Draining hosts accept no new leases; existing ones finish.
- Table persisted to
state.jsonafter every change; loaded at start so a crossbar restart does not reshuffle sessions.
6. Health
- Poller, every
poll_interval(60 s):GET /healththenGET /modelsper host. Record loaded models. Only thenGET /slots?model=Xfor loaded models only — probing an unloaded model makes llama-server load it. Per-model free-slot count from/slots(or, if/slotsis disabled on a host, assumeparallelminus our own in-flight count). - Passive: a connection error or 5xx on a proxied request marks the host unhealthy immediately and re-leases the route on the client's retry. Recovery only through the poller (two consecutive good polls).
- Concurrency: per
(host, model)token bucket of sizeparallel. Requests beyond it wait in a bounded FIFO (queue_max, default 8; 503 beyond that). This is where the unified-KV overflow is prevented.
7. Admin API (tailnet only, same listener, prefix /_crossbar)
GET /_crossbar/hosts— health, loaded models, free slots, in-flight.GET /_crossbar/routes— the lease table.POST /_crossbar/routes/{route}{ "host": "titan", "pin": true }— pin;{ "release": true }— drop the lease (next request re-chooses).POST /_crossbar/hosts/{host}{ "drain": true|false }.GET /_crossbar/metrics— Prometheus: requests, queue wait, lease moves, host health, tokens/s from llama-server'stimingswhen present. Scrape it from the fleet Prometheus on orion.
8. Config (crossbar.yaml)
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
poll_interval: 60s
lease_idle: 30m
queue_max: 8
hosts:
straylight:
base_url: http://straylight.<tailnet>:11434
weight: 1.0
models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
titan:
base_url: http://titan.<tailnet>:8081
weight: 2.0 # ~2× straylight decode
models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
dixie:
base_url: http://dixie.<tailnet>:11434
weight: 0.8
models: { ornith-1.5-9b-uncensored: {parallel: 6} }
routes: # optional defaults / pins
opencode-a: { default_model: ornith-1.5-35b-a3b }
opencode-b: { default_model: ornith-1.5-35b-a3b }
paper: { default_model: qwen3.8-27b-uncensored, pin: titan }
hermes-straylight: {}
hermes-talos: {}
hermes-titan: { pin: titan } # titan's models are the titan agent's first
Client side, no code changes:
- OpenCode: project-local
opencode.jsonprovider withoptions.baseURL: http://crossbar:7777/opencode-a/v1(global + project configs merge, so each project carries its own route). - Hermes:
custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1(and thedelegation/auxiliaryblocks that point at a router today). - tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (
plain_http_to_sinkfor that host only).
9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (
tailscale whoison the peer address) and restrict routes to peers, sohermes-taloscan only be used from talos. - Admin API: same listener, same trust. Consider a separate
admin_listenon localhost if crossbar runs on a shared host. - No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and timings only.
10. Milestones
- v0 (a day): config, static routes → host mapping, reverse proxy with
streaming,
/health+/modelspoller, passive health,/_crossbar/hosts. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins. - v1: lease table with persistence,
choose()by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. - v2: wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ 45 s, else fall through); optional conversation fingerprint (system prompt + first user message) for per-conversation stickiness inside one route; Tailscale-identity route restriction.
11. Testing
httptestfake llama-server:/health,/models,/slots, streaming/v1/chat/completionswith configurable latency and failure injection.- Table tests for
choose(), lease expiry, drain, passive-health re-lease. - One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server
timings.cache_norprompt_nsmall after the first turn).
12. Open questions
- Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static Go binary + systemd unit; arm64 cross-compile is free. Candidate: hyperborea, in its own unit, outside the collector's disk budget.
/slotsis disabled by default on recent llama-server builds; enable--slotson the three routers or rely on our own in-flight counts.- Should
default_modelrewrite a client'smodelfield? Proposal: no — proxy what the client sent; only use the default when the body has none. - Does the titan agent's reservation (2026-09-21) stand as a pin, or does titan become a general pool member now that health is tracked? Kyle said titan could join the pool (2026-09-25); the pin above keeps its own agent first without excluding others.