189 lines
9.0 KiB
Markdown
189 lines
9.0 KiB
Markdown
# crossbar — an affinity router for the fleet's llama-servers
|
||
|
||
**Status:** plan, 2026-09-25. Nothing built. Go.
|
||
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
|
||
|
||
## 1. Problem
|
||
|
||
Three hosts run `llama-server` in router mode with the *same model ids*
|
||
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
|
||
the paper writer) each hard-code one host. Consequences seen this week:
|
||
|
||
- a host's prompt cache is per process; an agent session sits at 100–230k
|
||
tokens with ~100 % cache hits. Any move between hosts re-prefills the whole
|
||
context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
|
||
- two turns landing on one 35B at once overflowed its unified KV
|
||
("Context size has been exceeded"); nothing in the path knows the slot count;
|
||
- titan sleeps; clients pointed at it fail instead of falling back;
|
||
- rebuilding or restarting a host takes its clients down with it.
|
||
|
||
A crossbar switch connects each caller to a line and *holds the connection*.
|
||
That is the design: **affinity first, balancing only for new sessions.**
|
||
|
||
## 2. Non-goals
|
||
|
||
- No token counting, prompt rewriting, batching or splitting across hosts.
|
||
The hosts are bandwidth-bound; there is nothing to gain cross-host.
|
||
- No model management. Loading/unloading stays with each host's llama-server
|
||
router (`--models-max`). crossbar never sends a request that would make a
|
||
host load a model it does not have resident.
|
||
- No auth in v0 (tailnet-only bind). See §9.
|
||
|
||
## 3. Terminology
|
||
|
||
- **host** — one llama-server router endpoint (`base_url`, list of model ids
|
||
it can serve, per-model `parallel`, a speed `weight`).
|
||
- **route** — a client identity = the first path segment of the request URL
|
||
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
|
||
base URL, so this needs no client support beyond "configurable base URL".
|
||
- **lease** — `(route, model) → host`, with `created`, `last_used`,
|
||
`state ∈ {active, pinned, draining}`.
|
||
|
||
## 4. Request flow
|
||
|
||
```
|
||
client ── /{route}/v1/chat/completions ──▶ crossbar
|
||
1. route := first path segment; strip it
|
||
2. model := body.model (JSON peek; fall back to route default)
|
||
3. lease := table[(route, model)]
|
||
hit & host healthy & model resident → use it
|
||
miss | host unhealthy → choose(host) ; write lease
|
||
4. acquire one concurrency token for (host, model) [bounded queue]
|
||
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
|
||
flush on every chunk; passive health on connect error / 5xx
|
||
6. release token; lease.last_used = now
|
||
```
|
||
|
||
`choose(host)`: among hosts that are healthy **and list the model as loaded**
|
||
(`/models`), take the one with the most free slots × `weight`; tie → lowest
|
||
current queue depth. If no host has it loaded, take a healthy host that *can*
|
||
serve it (config) and accept the load; log that decision.
|
||
|
||
Pass-through per route, all answered from the leased host (or from config if
|
||
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
|
||
`/v1/completions`. Everything else 404.
|
||
|
||
## 5. Lease rules
|
||
|
||
- A lease is **sticky**. It moves only when: the host fails health, the route
|
||
is idle longer than `lease_idle` (default 30 min), or an operator pins/
|
||
releases it. A faster host coming back online does *not* move an active
|
||
lease — the re-prefill is the cost we are avoiding.
|
||
- **Pinned** leases never move automatically ("project A goes to titan right
|
||
now"). **Draining** hosts accept no new leases; existing ones finish.
|
||
- Table persisted to `state.json` after every change; loaded at start so a
|
||
crossbar restart does not reshuffle sessions.
|
||
|
||
## 6. Health
|
||
|
||
- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per
|
||
host. Record loaded models. Only then `GET /slots?model=X` **for loaded
|
||
models only** — probing an unloaded model makes llama-server load it.
|
||
Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a
|
||
host, assume `parallel` minus our own in-flight count).
|
||
- Passive: a connection error or 5xx on a proxied request marks the host
|
||
unhealthy immediately and re-leases the route on the client's retry.
|
||
Recovery only through the poller (two consecutive good polls).
|
||
- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests
|
||
beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that).
|
||
This is where the unified-KV overflow is prevented.
|
||
|
||
## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`)
|
||
|
||
- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight.
|
||
- `GET /_crossbar/routes` — the lease table.
|
||
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
|
||
`{ "release": true }` — drop the lease (next request re-chooses).
|
||
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
|
||
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
|
||
host health, tokens/s from llama-server's `timings` when present. Scrape it
|
||
from the fleet Prometheus on orion.
|
||
|
||
## 8. Config (`crossbar.yaml`)
|
||
|
||
```yaml
|
||
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
|
||
poll_interval: 60s
|
||
lease_idle: 30m
|
||
queue_max: 8
|
||
hosts:
|
||
straylight:
|
||
base_url: http://straylight.<tailnet>:11434
|
||
weight: 1.0
|
||
models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
|
||
titan:
|
||
base_url: http://titan.<tailnet>:8081
|
||
weight: 2.0 # ~2× straylight decode
|
||
models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
|
||
dixie:
|
||
base_url: http://dixie.<tailnet>:11434
|
||
weight: 0.8
|
||
models: { ornith-1.5-9b-uncensored: {parallel: 6} }
|
||
routes: # optional defaults / pins
|
||
opencode-a: { default_model: ornith-1.5-35b-a3b }
|
||
opencode-b: { default_model: ornith-1.5-35b-a3b }
|
||
paper: { default_model: qwen3.8-27b-uncensored, pin: titan }
|
||
hermes-straylight: {}
|
||
hermes-talos: {}
|
||
hermes-titan: { pin: titan } # titan's models are the titan agent's first
|
||
```
|
||
|
||
Client side, no code changes:
|
||
|
||
- OpenCode: project-local `opencode.json` provider with
|
||
`options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project
|
||
configs merge, so each project carries its own route).
|
||
- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1`
|
||
(and the `delegation` / `auxiliary` blocks that point at a router today).
|
||
- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
|
||
the llama routers already have (`plain_http_to_sink` for that host only).
|
||
|
||
## 9. Security notes
|
||
|
||
- Bind to the tailnet address only. v0 relies on the tailnet for
|
||
authentication; every route is reachable by every tailnet peer, which is the
|
||
same exposure the llama-servers have today. v1 option: read Tailscale
|
||
identity (`tailscale whois` on the peer address) and restrict routes to
|
||
peers, so `hermes-talos` can only be used from talos.
|
||
- Admin API: same listener, same trust. Consider a separate `admin_listen` on
|
||
localhost if crossbar runs on a shared host.
|
||
- No secrets in config; llama-servers take no keys.
|
||
- Request bodies are proxied, never logged. Metrics carry counts and
|
||
timings only.
|
||
|
||
## 10. Milestones
|
||
|
||
- **v0 (a day):** config, static routes → host mapping, reverse proxy with
|
||
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
|
||
Replaces the hand-maintained provider lists. No leases yet: each route has
|
||
a fixed host list in preference order; first healthy wins.
|
||
- **v1:** lease table with persistence, `choose()` by free slots × weight,
|
||
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
|
||
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
|
||
45 s, else fall through); optional conversation fingerprint (system prompt +
|
||
first user message) for per-conversation stickiness inside one route;
|
||
Tailscale-identity route restriction.
|
||
|
||
## 11. Testing
|
||
|
||
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
|
||
`/v1/chat/completions` with configurable latency and failure injection.
|
||
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
|
||
- One integration script against the real fleet: one long conversation,
|
||
assert every turn hits the same host (llama-server `timings.cache_n` or
|
||
`prompt_n` small after the first turn).
|
||
|
||
## 12. Open questions
|
||
|
||
1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static
|
||
Go binary + systemd unit; arm64 cross-compile is free. Candidate:
|
||
hyperborea, in its own unit, outside the collector's disk budget.
|
||
2. `/slots` is disabled by default on recent llama-server builds; enable
|
||
`--slots` on the three routers or rely on our own in-flight counts.
|
||
3. Should `default_model` rewrite a client's `model` field? Proposal: no —
|
||
proxy what the client sent; only use the default when the body has none.
|
||
4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does
|
||
titan become a general pool member now that health is tracked? Kyle said
|
||
titan could join the pool (2026-09-25); the pin above keeps its own agent
|
||
first without excluding others.
|