Files
crossbar/PLAN.md
T

239 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# crossbar — an affinity router for the fleet's llama-servers
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting, Kyle). Nothing built. Go.
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
## 1. Problem
Three hosts run `llama-server` in router mode with the *same model ids*
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k
tokens with ~100 % cache hits. Any move between hosts re-prefills the whole
context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV
("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and *holds the connection*.
That is the design: **affinity first, balancing only for new sessions.**
## 2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts.
The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (`--models-max`). crossbar never sends a request that would make a
host load a model it does not have resident.
- No auth in v0 (tailnet-only bind). See §9.
## 3. Terminology
- **host** — one llama-server router endpoint (`base_url`, list of model ids
it can serve, per-model `parallel`, a speed `weight`).
- **route** — a client identity = the first path segment of the request URL
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
base URL, so this needs no client support beyond "configurable base URL".
- **lease** — `(route, model) → host`, with `created`, `last_used`,
`state ∈ {active, pinned, draining}`.
## 4. Request flow
```
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
3. lease := table[(route, model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
```
`choose(host)`: among hosts that are healthy **and list the model as loaded**
(`/models`), take the one with the most free slots × `weight`; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that *can*
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
`/v1/completions`. Everything else 404.
## 5. Lease rules
- A lease is **sticky**. It moves only when: the host fails health, the route
is idle longer than `lease_idle` (default 30 min), or an operator pins/
releases it. A faster host coming back online does *not* move an active
lease — the re-prefill is the cost we are avoiding.
- **Pinned** leases never move automatically ("project A goes to titan right
now"). **Draining** hosts accept no new leases; existing ones finish.
- Table lives in memory and is written through to SQLite (§7a) on every
change; loaded at start so a crossbar restart does not reshuffle sessions.
## 6. Health
- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per
host. Record loaded models. Only then `GET /slots?model=X` **for loaded
models only** — probing an unloaded model makes llama-server load it.
Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a
host, assume `parallel` minus our own in-flight count).
- Passive: a connection error or 5xx on a proxied request marks the host
unhealthy immediately and re-leases the route on the client's retry.
Recovery only through the poller (two consecutive good polls).
- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests
beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that).
This is where the unified-KV overflow is prevented.
## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`)
- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight.
- `GET /_crossbar/routes` — the lease table.
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
`{ "release": true }` — drop the lease (next request re-chooses).
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
- `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups
from §7a (JSON; `Accept: text/plain` gives a table).
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
host health, tokens/s from llama-server's `timings` when present. Scrape it
from the fleet Prometheus on orion.
## 7a. State store and accounting (SQLite)
One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the
accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64
static build stays a plain `go build`. Single writer goroutine fed by a
channel; readers use their own connection. Volume is a few rows per request,
so nothing here is performance-sensitive.
```sql
CREATE TABLE leases ( -- current table, one row per (route, model)
route TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
created INTEGER, last_used INTEGER, PRIMARY KEY (route, model));
CREATE TABLE lease_events ( -- why sessions moved
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
reason TEXT); -- new|unhealthy|idle|pin|release|drain
CREATE TABLE requests ( -- one row per proxied completion
id INTEGER PRIMARY KEY, route TEXT, model TEXT, host TEXT,
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
status INTEGER, streamed INTEGER,
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
err TEXT);
CREATE TABLE host_health ( -- poller observations, for uptime accounting
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
```
Token and cache figures come from the upstream response when llama-server
provides them: `usage` on non-streaming replies, and the final SSE chunk's
`usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on
streamed replies. To see that chunk the proxy tees the response body through
a small SSE line scanner; it never buffers or alters the stream. When the
fields are absent, wall-clock columns are still recorded.
What this answers: per route (session/project), per model, per host — number
of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio
(`cached_tokens / prompt_tokens`, the direct measure of whether affinity is
working), queue wait, error rate, and per-host uptime. `/_crossbar/usage`
exposes the rollups; a nightly job prunes `requests` older than `retention`
(default 180 d) into a `requests_daily` rollup so the file stays small.
Body contents are never stored — only counts and timings.
## 8. Config (`crossbar.yaml`)
```yaml
listen: "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
db: /var/lib/crossbar/crossbar.db
retention: 180d
poll_interval: 60s
lease_idle: 30m
queue_max: 8
hosts:
straylight:
base_url: http://straylight.<tailnet>:11434
weight: 1.0
models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
titan:
base_url: http://titan.<tailnet>:8081
weight: 2.0 # ~2× straylight decode
models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
dixie:
base_url: http://dixie.<tailnet>:11434
weight: 0.8
models: { ornith-1.5-9b-uncensored: {parallel: 6} }
routes: # optional defaults / pins
opencode-a: { default_model: ornith-1.5-35b-a3b }
opencode-b: { default_model: ornith-1.5-35b-a3b }
paper: { default_model: qwen3.8-27b-uncensored, pin: titan }
hermes-straylight: {}
hermes-talos: {}
hermes-titan: { pin: titan } # titan's models are the titan agent's first
```
Client side, no code changes:
- OpenCode: project-local `opencode.json` provider with
`options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project
configs merge, so each project carries its own route).
- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1`
(and the `delegation` / `auxiliary` blocks that point at a router today).
- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (`plain_http_to_sink` for that host only).
## 9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (`tailscale whois` on the peer address) and restrict routes to
peers, so `hermes-talos` can only be used from talos.
- Admin API: same listener, same trust. Consider a separate `admin_listen` on
localhost if crossbar runs on a shared host.
- No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and
timings only.
## 10. Milestones
- **v0 (a day):** config, static routes → host mapping, reverse proxy with
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
Replaces the hand-maintained provider lists. No leases yet: each route has
a fixed host list in preference order; first healthy wins.
- **v1:** SQLite state + accounting (§7a), lease table, `choose()` by free slots × weight,
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
45 s, else fall through); optional conversation fingerprint (system prompt +
first user message) for per-conversation stickiness inside one route;
Tailscale-identity route restriction.
## 11. Testing
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
`/v1/chat/completions` with configurable latency and failure injection.
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease.
- SSE tee scanner: fixture streams with and without a final `usage` chunk;
assert the client receives the bytes unchanged and the row is recorded.
- One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server `timings.cache_n` or
`prompt_n` small after the first turn).
## 12. Open questions
1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static
Go binary + systemd unit; arm64 cross-compile is free. Candidate:
hyperborea, in its own unit, outside the collector's disk budget.
2. `/slots` is disabled by default on recent llama-server builds; enable
`--slots` on the three routers or rely on our own in-flight counts.
3. Should `default_model` rewrite a client's `model` field? Proposal: no —
proxy what the client sent; only use the default when the body has none.
4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does
titan become a general pool member now that health is tracked? Kyle said
titan could join the pool (2026-09-25); the pin above keeps its own agent
first without excluding others.