commit 2c6bd7b2ce2f66a436256c5a8ceece803c124f20 Author: Kyle Isom Date: Fri Sep 25 00:41:16 2026 -0700 PLAN: crossbar, an affinity router for the fleet llama-servers Co-Authored-By: Claude Fable 5.1 diff --git a/PLAN.md b/PLAN.md new file mode 100644 index 0000000..7391daa --- /dev/null +++ b/PLAN.md @@ -0,0 +1,188 @@ +# crossbar — an affinity router for the fleet's llama-servers + +**Status:** plan, 2026-09-25. Nothing built. Go. +**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure). + +## 1. Problem + +Three hosts run `llama-server` in router mode with the *same model ids* +(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents, +the paper writer) each hard-code one host. Consequences seen this week: + +- a host's prompt cache is per process; an agent session sits at 100–230k + tokens with ~100 % cache hits. Any move between hosts re-prefills the whole + context (≈ 7–8 min at straylight's ~500 t/s prompt processing); +- two turns landing on one 35B at once overflowed its unified KV + ("Context size has been exceeded"); nothing in the path knows the slot count; +- titan sleeps; clients pointed at it fail instead of falling back; +- rebuilding or restarting a host takes its clients down with it. + +A crossbar switch connects each caller to a line and *holds the connection*. +That is the design: **affinity first, balancing only for new sessions.** + +## 2. Non-goals + +- No token counting, prompt rewriting, batching or splitting across hosts. + The hosts are bandwidth-bound; there is nothing to gain cross-host. +- No model management. Loading/unloading stays with each host's llama-server + router (`--models-max`). crossbar never sends a request that would make a + host load a model it does not have resident. +- No auth in v0 (tailnet-only bind). See §9. + +## 3. Terminology + +- **host** — one llama-server router endpoint (`base_url`, list of model ids + it can serve, per-model `parallel`, a speed `weight`). +- **route** — a client identity = the first path segment of the request URL + (`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a + base URL, so this needs no client support beyond "configurable base URL". +- **lease** — `(route, model) → host`, with `created`, `last_used`, + `state ∈ {active, pinned, draining}`. + +## 4. Request flow + +``` +client ── /{route}/v1/chat/completions ──▶ crossbar + 1. route := first path segment; strip it + 2. model := body.model (JSON peek; fall back to route default) + 3. lease := table[(route, model)] + hit & host healthy & model resident → use it + miss | host unhealthy → choose(host) ; write lease + 4. acquire one concurrency token for (host, model) [bounded queue] + 5. reverse-proxy to host.base_url + "/v1/..." ; stream through, + flush on every chunk; passive health on connect error / 5xx + 6. release token; lease.last_used = now +``` + +`choose(host)`: among hosts that are healthy **and list the model as loaded** +(`/models`), take the one with the most free slots × `weight`; tie → lowest +current queue depth. If no host has it loaded, take a healthy host that *can* +serve it (config) and accept the load; log that decision. + +Pass-through per route, all answered from the leased host (or from config if +no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`, +`/v1/completions`. Everything else 404. + +## 5. Lease rules + +- A lease is **sticky**. It moves only when: the host fails health, the route + is idle longer than `lease_idle` (default 30 min), or an operator pins/ + releases it. A faster host coming back online does *not* move an active + lease — the re-prefill is the cost we are avoiding. +- **Pinned** leases never move automatically ("project A goes to titan right + now"). **Draining** hosts accept no new leases; existing ones finish. +- Table persisted to `state.json` after every change; loaded at start so a + crossbar restart does not reshuffle sessions. + +## 6. Health + +- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per + host. Record loaded models. Only then `GET /slots?model=X` **for loaded + models only** — probing an unloaded model makes llama-server load it. + Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a + host, assume `parallel` minus our own in-flight count). +- Passive: a connection error or 5xx on a proxied request marks the host + unhealthy immediately and re-leases the route on the client's retry. + Recovery only through the poller (two consecutive good polls). +- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests + beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that). + This is where the unified-KV overflow is prevented. + +## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`) + +- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight. +- `GET /_crossbar/routes` — the lease table. +- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin; + `{ "release": true }` — drop the lease (next request re-chooses). +- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`. +- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves, + host health, tokens/s from llama-server's `timings` when present. Scrape it + from the fleet Prometheus on orion. + +## 8. Config (`crossbar.yaml`) + +```yaml +listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0 +poll_interval: 60s +lease_idle: 30m +queue_max: 8 +hosts: + straylight: + base_url: http://straylight.:11434 + weight: 1.0 + models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} } + titan: + base_url: http://titan.:8081 + weight: 2.0 # ~2× straylight decode + models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} } + dixie: + base_url: http://dixie.:11434 + weight: 0.8 + models: { ornith-1.5-9b-uncensored: {parallel: 6} } +routes: # optional defaults / pins + opencode-a: { default_model: ornith-1.5-35b-a3b } + opencode-b: { default_model: ornith-1.5-35b-a3b } + paper: { default_model: qwen3.8-27b-uncensored, pin: titan } + hermes-straylight: {} + hermes-talos: {} + hermes-titan: { pin: titan } # titan's models are the titan agent's first +``` + +Client side, no code changes: + +- OpenCode: project-local `opencode.json` provider with + `options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project + configs merge, so each project carries its own route). +- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-/v1` + (and the `delegation` / `auxiliary` blocks that point at a router today). +- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry + the llama routers already have (`plain_http_to_sink` for that host only). + +## 9. Security notes + +- Bind to the tailnet address only. v0 relies on the tailnet for + authentication; every route is reachable by every tailnet peer, which is the + same exposure the llama-servers have today. v1 option: read Tailscale + identity (`tailscale whois` on the peer address) and restrict routes to + peers, so `hermes-talos` can only be used from talos. +- Admin API: same listener, same trust. Consider a separate `admin_listen` on + localhost if crossbar runs on a shared host. +- No secrets in config; llama-servers take no keys. +- Request bodies are proxied, never logged. Metrics carry counts and + timings only. + +## 10. Milestones + +- **v0 (a day):** config, static routes → host mapping, reverse proxy with + streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`. + Replaces the hand-maintained provider lists. No leases yet: each route has + a fixed host list in preference order; first healthy wins. +- **v1:** lease table with persistence, `choose()` by free slots × weight, + per-(host, model) concurrency + bounded queue, pin/release/drain, metrics. +- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ + 45 s, else fall through); optional conversation fingerprint (system prompt + + first user message) for per-conversation stickiness inside one route; + Tailscale-identity route restriction. + +## 11. Testing + +- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming + `/v1/chat/completions` with configurable latency and failure injection. +- Table tests for `choose()`, lease expiry, drain, passive-health re-lease. +- One integration script against the real fleet: one long conversation, + assert every turn hits the same host (llama-server `timings.cache_n` or + `prompt_n` small after the first turn). + +## 12. Open questions + +1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static + Go binary + systemd unit; arm64 cross-compile is free. Candidate: + hyperborea, in its own unit, outside the collector's disk budget. +2. `/slots` is disabled by default on recent llama-server builds; enable + `--slots` on the three routers or rely on our own in-flight counts. +3. Should `default_model` rewrite a client's `model` field? Proposal: no — + proxy what the client sent; only use the default when the body has none. +4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does + titan become a general pool member now that health is tracked? Kyle said + titan could join the pool (2026-09-25); the pin above keeps its own agent + first without excluding others.