Files
crossbar/PLAN.md
T
2026-09-25 00:49:30 -07:00

296 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# crossbar — an affinity router for the fleet's llama-servers
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go.
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
## 1. Problem
Three hosts run `llama-server` in router mode with the *same model ids*
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
the paper writer) each hard-code one host. Consequences seen this week:
- a host's prompt cache is per process; an agent session sits at 100–230k
tokens with ~100 % cache hits. Any move between hosts re-prefills the whole
context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
- two turns landing on one 35B at once overflowed its unified KV
("Context size has been exceeded"); nothing in the path knows the slot count;
- titan sleeps; clients pointed at it fail instead of falling back;
- rebuilding or restarting a host takes its clients down with it.
A crossbar switch connects each caller to a line and *holds the connection*.
That is the design: **affinity first, balancing only for new sessions.**
## 2. Non-goals
- No token counting, prompt rewriting, batching or splitting across hosts.
The hosts are bandwidth-bound; there is nothing to gain cross-host.
- No model management. Loading/unloading stays with each host's llama-server
router (`--models-max`). crossbar never sends a request that would make a
host load a model it does not have resident.
- No auth in v0 (tailnet-only bind). See §9.
## 3. Terminology
- **host** — one llama-server router endpoint (`base_url`, list of model ids
it can serve, per-model `parallel`, a speed `weight`).
- **route** — a client identity = the first path segment of the request URL
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
base URL, so this needs no client support beyond "configurable base URL".
- **fingerprint** — a per-conversation key inside a route: SHA-256 of the
request's system prompt plus its first user message (first 4 KiB of each).
Stable across a conversation's turns, different between conversations. Empty
when the body has no user message (probes, title generation).
- **lease** — `(route, fingerprint, model) → host`, with `created`, `last_used`,
`state ∈ {active, pinned, draining}`. Pins are per route; a pinned route
pins all its fingerprints.
## 4. Request flow
```
client ── /{route}/v1/chat/completions ──▶ crossbar
1. route := first path segment; strip it
2. model := body.model (JSON peek; fall back to route default)
2b. fp := fingerprint(body) ("" if none)
3. lease := table[(route, fp, model)] → fallback table[(route, "", model)]
hit & host healthy & model resident → use it
miss | host unhealthy → choose(host) ; write lease
4. acquire one concurrency token for (host, model) [bounded queue]
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
flush on every chunk; passive health on connect error / 5xx
6. release token; lease.last_used = now
```
`choose(host)`: among hosts that are healthy **and list the model as loaded**
(`/models`), take the one with the most free slots × `weight`; tie → lowest
current queue depth. If no host has it loaded, take a healthy host that *can*
serve it (config) and accept the load; log that decision.
Pass-through per route, all answered from the leased host (or from config if
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
`/v1/completions`. Everything else 404.
## 4a. Why two keys (Kyle's main use case: many OpenCode instances on one box)
Per-instance identity comes from the route; per-session identity comes from
the fingerprint. OpenCode's system prompt is per project and its first user
message is per session, so `(route, fp)` separates sessions inside one
instance without any client support. N sessions then spread across hosts by
free slots × weight *at start* and are held there; the per-(host, model) queue
(§6) is what stops N from oversubscribing any one host.
Client launcher (OpenCode config substitutes `{env:VAR}` in values — verified
in the installed 1.15.10 bundle: `/\{env:([^}]+)\}/g` → `process.env[V]`,
missing → empty string, so always set the variable; as a fallback the same
value can go through `options.headers["X-Crossbar-Route"]`, which crossbar
also accepts):
```jsonc
// ~/.config/opencode/opencode.json (one block for every project)
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
"options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" },
"models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } }
```
```sh
# oc: one route per instance
CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@"
```
## 4b. Context sizes: learned, not configured
The poller records `n_ctx`, `n_parallel` (→ per-slot context with unified KV)
from `/props` per host and model into `host_health`; `/props` is passed
through per route from the leased host so clients see the real number. Config
may override (`ctx = N` under a model) but normally does not. v2 uses it as a
guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased
host's per-slot context is re-leased to a host where it fits, or answered
400 with a clear message instead of the upstream "Context size has been
exceeded".
## 5. Lease rules
- A lease is **sticky**. It moves only when: the host fails health, the route
is idle longer than `lease_idle` (default 30 min), or an operator pins/
releases it. A faster host coming back online does *not* move an active
lease — the re-prefill is the cost we are avoiding.
- **Pinned** leases never move automatically ("project A goes to titan right
now"). **Draining** hosts accept no new leases; existing ones finish.
- Table lives in memory and is written through to SQLite (§7a) on every
change; loaded at start so a crossbar restart does not reshuffle sessions.
## 6. Health
- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per
host. Record loaded models. Only then `GET /slots?model=X` **for loaded
models only** — probing an unloaded model makes llama-server load it.
Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a
host, assume `parallel` minus our own in-flight count).
- Passive: a connection error or 5xx on a proxied request marks the host
unhealthy immediately and re-leases the route on the client's retry.
Recovery only through the poller (two consecutive good polls).
- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests
beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that).
This is where the unified-KV overflow is prevented.
## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`)
- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight.
- `GET /_crossbar/routes` — the lease table.
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
`{ "release": true }` — drop the lease (next request re-chooses).
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
- `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups
from §7a (JSON; `Accept: text/plain` gives a table).
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
host health, tokens/s from llama-server's `timings` when present. Scrape it
from the fleet Prometheus on orion.
## 7a. State store and accounting (SQLite)
One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the
accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64
static build stays a plain `go build`. Single writer goroutine fed by a
channel; readers use their own connection. Volume is a few rows per request,
so nothing here is performance-sensitive.
```sql
CREATE TABLE leases ( -- current table, one row per (route, model)
route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model));
CREATE TABLE lease_events ( -- why sessions moved
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
reason TEXT); -- new|unhealthy|idle|pin|release|drain
CREATE TABLE requests ( -- one row per proxied completion
id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT,
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
status INTEGER, streamed INTEGER,
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
err TEXT);
CREATE TABLE host_health ( -- poller observations, for uptime accounting
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
```
Token and cache figures come from the upstream response when llama-server
provides them: `usage` on non-streaming replies, and the final SSE chunk's
`usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on
streamed replies. To see that chunk the proxy tees the response body through
a small SSE line scanner; it never buffers or alters the stream. When the
fields are absent, wall-clock columns are still recorded.
What this answers: per route (session/project), per model, per host — number
of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio
(`cached_tokens / prompt_tokens`, the direct measure of whether affinity is
working), queue wait, error rate, and per-host uptime. `/_crossbar/usage`
exposes the rollups; a nightly job prunes `requests` older than `retention`
(default 180 d) into a `requests_daily` rollup so the file stays small.
Body contents are never stored — only counts and timings.
## 8. Config (`crossbar.toml`)
TOML (`github.com/BurntSushi/toml`): no indentation semantics, no implicit
type coercion, and it is what Kyle's Rust projects already use. A NixOS module
can generate it with `pkgs.formats.toml`.
```toml
listen = "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
db = "/var/lib/crossbar/crossbar.db"
poll_interval = "60s"
lease_idle = "30m"
queue_max = 8
retention = "180d"
[hosts.straylight]
base_url = "http://straylight.<tailnet>:11434"
weight = 1.0
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "ornith-1.5-9b-uncensored" = { parallel = 6 } }
[hosts.titan]
base_url = "http://titan.<tailnet>:8081"
weight = 2.0 # ~2x straylight decode
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "laguna-s-2.1" = { parallel = 2 } }
[hosts.dixie]
base_url = "http://dixie.<tailnet>:11434"
weight = 0.8
models = { "ornith-1.5-9b-uncensored" = { parallel = 6 } }
# optional per-route defaults / pins
[routes.opencode-a]
default_model = "ornith-1.5-35b-a3b"
[routes.opencode-b]
default_model = "ornith-1.5-35b-a3b"
[routes.paper]
default_model = "qwen3.8-27b-uncensored"
pin = "titan"
[routes.hermes-straylight]
[routes.hermes-talos]
[routes.hermes-titan]
pin = "titan" # titan's models are the titan agent's first
```
Client side, no code changes:
- OpenCode: project-local `opencode.json` provider with
`options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project
configs merge, so each project carries its own route).
- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1`
(and the `delegation` / `auxiliary` blocks that point at a router today).
- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
the llama routers already have (`plain_http_to_sink` for that host only).
## 9. Security notes
- Bind to the tailnet address only. v0 relies on the tailnet for
authentication; every route is reachable by every tailnet peer, which is the
same exposure the llama-servers have today. v1 option: read Tailscale
identity (`tailscale whois` on the peer address) and restrict routes to
peers, so `hermes-talos` can only be used from talos.
- Admin API: same listener, same trust. Consider a separate `admin_listen` on
localhost if crossbar runs on a shared host.
- No secrets in config; llama-servers take no keys.
- Request bodies are proxied, never logged. Metrics carry counts and
timings only.
## 10. Milestones
- **v0 (a day):** TOML config, static routes → host mapping, reverse proxy with
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
Replaces the hand-maintained provider lists. No leases yet: each route has
a fixed host list in preference order; first healthy wins.
- **v1:** SQLite state + accounting (§7a), lease table keyed by
`(route, fingerprint, model)`, header route override, `choose()` by free slots × weight,
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
45 s, else fall through); context-size guard (§4b); Tailscale-identity
route restriction.
## 11. Testing
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
`/v1/chat/completions` with configurable latency and failure injection.
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease,
fingerprint stability (same conversation → same key; different first user
message → different key; no user message → empty key).
- SSE tee scanner: fixture streams with and without a final `usage` chunk;
assert the client receives the bytes unchanged and the row is recorded.
- One integration script against the real fleet: one long conversation,
assert every turn hits the same host (llama-server `timings.cache_n` or
`prompt_n` small after the first turn).
## 12. Open questions
1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static
Go binary + systemd unit; arm64 cross-compile is free. Candidate:
hyperborea, in its own unit, outside the collector's disk budget.
2. `/slots` is disabled by default on recent llama-server builds; enable
`--slots` on the three routers or rely on our own in-flight counts.
3. Should `default_model` rewrite a client's `model` field? Proposal: no —
proxy what the client sent; only use the default when the body has none.
4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does
titan become a general pool member now that health is tracked? Kyle said
titan could join the pool (2026-09-25); the pin above keeps its own agent
first without excluding others.