Files
crossbar/PLAN.md
T

189 lines
9.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# crossbar — an affinity router for the fleet's llama-servers
**Status:** plan, 2026-09-25. Nothing built. Go.
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
## 1. Problem
Three hosts run `llama-server` in router mode with the *same model ids*
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k
tokens with ~100 % cache hits. Any move between hosts re-prefills the whole
context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV
("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and *holds the connection*.
That is the design: **affinity first, balancing only for new sessions.**
## 2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts.
The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (`--models-max`). crossbar never sends a request that would make a
host load a model it does not have resident.
- No auth in v0 (tailnet-only bind). See §9.
## 3. Terminology
- **host** — one llama-server router endpoint (`base_url`, list of model ids
it can serve, per-model `parallel`, a speed `weight`).
- **route** — a client identity = the first path segment of the request URL
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
base URL, so this needs no client support beyond "configurable base URL".
- **lease** — `(route, model) → host`, with `created`, `last_used`,
`state ∈ {active, pinned, draining}`.
## 4. Request flow
```
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
3. lease := table[(route, model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
```
`choose(host)`: among hosts that are healthy **and list the model as loaded**
(`/models`), take the one with the most free slots × `weight`; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that *can*
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
`/v1/completions`. Everything else 404.
## 5. Lease rules
- A lease is **sticky**. It moves only when: the host fails health, the route
is idle longer than `lease_idle` (default 30 min), or an operator pins/
releases it. A faster host coming back online does *not* move an active
lease — the re-prefill is the cost we are avoiding.
- **Pinned** leases never move automatically ("project A goes to titan right
now"). **Draining** hosts accept no new leases; existing ones finish.
- Table persisted to `state.json` after every change; loaded at start so a
crossbar restart does not reshuffle sessions.
## 6. Health
- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per
host. Record loaded models. Only then `GET /slots?model=X` **for loaded
models only** — probing an unloaded model makes llama-server load it.
Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a
host, assume `parallel` minus our own in-flight count).
- Passive: a connection error or 5xx on a proxied request marks the host
unhealthy immediately and re-leases the route on the client's retry.
Recovery only through the poller (two consecutive good polls).
- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests
beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that).
This is where the unified-KV overflow is prevented.
## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`)
- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight.
- `GET /_crossbar/routes` — the lease table.
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
`{ "release": true }` — drop the lease (next request re-chooses).
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
host health, tokens/s from llama-server's `timings` when present. Scrape it
from the fleet Prometheus on orion.
## 8. Config (`crossbar.yaml`)
```yaml
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
poll_interval: 60s
lease_idle: 30m
queue_max: 8
hosts:
straylight:
base_url: http://straylight.<tailnet>:11434
weight: 1.0
models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
titan:
base_url: http://titan.<tailnet>:8081
weight: 2.0 # ~2× straylight decode
models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
dixie:
base_url: http://dixie.<tailnet>:11434
weight: 0.8
models: { ornith-1.5-9b-uncensored: {parallel: 6} }
routes: # optional defaults / pins
opencode-a: { default_model: ornith-1.5-35b-a3b }
opencode-b: { default_model: ornith-1.5-35b-a3b }
paper: { default_model: qwen3.8-27b-uncensored, pin: titan }
hermes-straylight: {}
hermes-talos: {}
hermes-titan: { pin: titan } # titan's models are the titan agent's first
```
Client side, no code changes:
- OpenCode: project-local `opencode.json` provider with
`options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project
configs merge, so each project carries its own route).
- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1`
(and the `delegation` / `auxiliary` blocks that point at a router today).
- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (`plain_http_to_sink` for that host only).
## 9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (`tailscale whois` on the peer address) and restrict routes to
peers, so `hermes-talos` can only be used from talos.
- Admin API: same listener, same trust. Consider a separate `admin_listen` on
localhost if crossbar runs on a shared host.
- No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and
timings only.
## 10. Milestones
- **v0 (a day):** config, static routes → host mapping, reverse proxy with
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
Replaces the hand-maintained provider lists. No leases yet: each route has
a fixed host list in preference order; first healthy wins.
- **v1:** lease table with persistence, `choose()` by free slots × weight,
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
45 s, else fall through); optional conversation fingerprint (system prompt +
first user message) for per-conversation stickiness inside one route;
Tailscale-identity route restriction.
## 11. Testing
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
`/v1/chat/completions` with configurable latency and failure injection.
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
- One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server `timings.cache_n` or
`prompt_n` small after the first turn).
## 12. Open questions
1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static
Go binary + systemd unit; arm64 cross-compile is free. Candidate:
hyperborea, in its own unit, outside the collector's disk budget.
2. `/slots` is disabled by default on recent llama-server builds; enable
`--slots` on the three routers or rely on our own in-flight counts.
3. Should `default_model` rewrite a client's `model` field? Proposal: no —
proxy what the client sent; only use the default when the body has none.
4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does
titan become a general pool member now that health is tracked? Kyle said
titan could join the pool (2026-09-25); the pin above keeps its own agent
first without excluding others.