PLAN: crossbar, an affinity router for the fleet llama-servers

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 00:41:16 -07:00
co-authored by Claude Fable 5.1
commit 2c6bd7b2ce
+188
View File
@@ -0,0 +1,188 @@
# crossbar — an affinity router for the fleet's llama-servers
**Status:** plan, 2026-09-25. Nothing built. Go.
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
## 1. Problem
Three hosts run `llama-server` in router mode with the *same model ids*
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k
tokens with ~100 % cache hits. Any move between hosts re-prefills the whole
context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV
("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and *holds the connection*.
That is the design: **affinity first, balancing only for new sessions.**
## 2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts.
The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (`--models-max`). crossbar never sends a request that would make a
host load a model it does not have resident.
- No auth in v0 (tailnet-only bind). See §9.
## 3. Terminology
- **host** — one llama-server router endpoint (`base_url`, list of model ids
it can serve, per-model `parallel`, a speed `weight`).
- **route** — a client identity = the first path segment of the request URL
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
base URL, so this needs no client support beyond "configurable base URL".
- **lease** — `(route, model) → host`, with `created`, `last_used`,
`state ∈ {active, pinned, draining}`.
## 4. Request flow
```
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
3. lease := table[(route, model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
```
`choose(host)`: among hosts that are healthy **and list the model as loaded**
(`/models`), take the one with the most free slots × `weight`; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that *can*
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
`/v1/completions`. Everything else 404.
## 5. Lease rules
- A lease is **sticky**. It moves only when: the host fails health, the route
is idle longer than `lease_idle` (default 30 min), or an operator pins/
releases it. A faster host coming back online does *not* move an active
lease — the re-prefill is the cost we are avoiding.
- **Pinned** leases never move automatically ("project A goes to titan right
now"). **Draining** hosts accept no new leases; existing ones finish.
- Table persisted to `state.json` after every change; loaded at start so a
crossbar restart does not reshuffle sessions.
## 6. Health
- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per
host. Record loaded models. Only then `GET /slots?model=X` **for loaded
models only** — probing an unloaded model makes llama-server load it.
Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a
host, assume `parallel` minus our own in-flight count).
- Passive: a connection error or 5xx on a proxied request marks the host
unhealthy immediately and re-leases the route on the client's retry.
Recovery only through the poller (two consecutive good polls).
- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests
beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that).
This is where the unified-KV overflow is prevented.
## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`)
- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight.
- `GET /_crossbar/routes` — the lease table.
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
`{ "release": true }` — drop the lease (next request re-chooses).
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
host health, tokens/s from llama-server's `timings` when present. Scrape it
from the fleet Prometheus on orion.
## 8. Config (`crossbar.yaml`)
```yaml
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
poll_interval: 60s
lease_idle: 30m
queue_max: 8
hosts:
straylight:
base_url: http://straylight.<tailnet>:11434
weight: 1.0
models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
titan:
base_url: http://titan.<tailnet>:8081
weight: 2.0 # ~2× straylight decode
models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
dixie:
base_url: http://dixie.<tailnet>:11434
weight: 0.8
models: { ornith-1.5-9b-uncensored: {parallel: 6} }
routes: # optional defaults / pins
opencode-a: { default_model: ornith-1.5-35b-a3b }
opencode-b: { default_model: ornith-1.5-35b-a3b }
paper: { default_model: qwen3.8-27b-uncensored, pin: titan }
hermes-straylight: {}
hermes-talos: {}
hermes-titan: { pin: titan } # titan's models are the titan agent's first
```
Client side, no code changes:
- OpenCode: project-local `opencode.json` provider with
`options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project
configs merge, so each project carries its own route).
- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1`
(and the `delegation` / `auxiliary` blocks that point at a router today).
- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (`plain_http_to_sink` for that host only).
## 9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (`tailscale whois` on the peer address) and restrict routes to
peers, so `hermes-talos` can only be used from talos.
- Admin API: same listener, same trust. Consider a separate `admin_listen` on
localhost if crossbar runs on a shared host.
- No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and
timings only.
## 10. Milestones
- **v0 (a day):** config, static routes → host mapping, reverse proxy with
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
Replaces the hand-maintained provider lists. No leases yet: each route has
a fixed host list in preference order; first healthy wins.
- **v1:** lease table with persistence, `choose()` by free slots × weight,
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
45 s, else fall through); optional conversation fingerprint (system prompt +
first user message) for per-conversation stickiness inside one route;
Tailscale-identity route restriction.
## 11. Testing
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
`/v1/chat/completions` with configurable latency and failure injection.
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
- One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server `timings.cache_n` or
`prompt_n` small after the first turn).
## 12. Open questions
1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static
Go binary + systemd unit; arm64 cross-compile is free. Candidate:
hyperborea, in its own unit, outside the collector's disk budget.
2. `/slots` is disabled by default on recent llama-server builds; enable
`--slots` on the three routers or rely on our own in-flight counts.
3. Should `default_model` rewrite a client's `model` field? Proposal: no —
proxy what the client sent; only use the default when the body has none.
4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does
titan become a general pool member now that health is tracked? Kyle said
titan could join the pool (2026-09-25); the pin above keeps its own agent
first without excluding others.