296 lines
14 KiB
Markdown
296 lines
14 KiB
Markdown
# crossbar — an affinity router for the fleet's llama-servers
|
||
|
||
**Status:** plan, 2026-09-25 (rev 2: SQLite state + accounting; rev 3: TOML config; rev 4: per-session leases, learned context sizes — Kyle). Nothing built. Go.
|
||
**Owner:** Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).
|
||
|
||
## 1. Problem
|
||
|
||
Three hosts run `llama-server` in router mode with the *same model ids*
|
||
(straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents,
|
||
the paper writer) each hard-code one host. Consequences seen this week:
|
||
|
||
- a host's prompt cache is per process; an agent session sits at 100–230k
|
||
tokens with ~100 % cache hits. Any move between hosts re-prefills the whole
|
||
context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
|
||
- two turns landing on one 35B at once overflowed its unified KV
|
||
("Context size has been exceeded"); nothing in the path knows the slot count;
|
||
- titan sleeps; clients pointed at it fail instead of falling back;
|
||
- rebuilding or restarting a host takes its clients down with it.
|
||
|
||
A crossbar switch connects each caller to a line and *holds the connection*.
|
||
That is the design: **affinity first, balancing only for new sessions.**
|
||
|
||
## 2. Non-goals
|
||
|
||
- No token counting, prompt rewriting, batching or splitting across hosts.
|
||
The hosts are bandwidth-bound; there is nothing to gain cross-host.
|
||
- No model management. Loading/unloading stays with each host's llama-server
|
||
router (`--models-max`). crossbar never sends a request that would make a
|
||
host load a model it does not have resident.
|
||
- No auth in v0 (tailnet-only bind). See §9.
|
||
|
||
## 3. Terminology
|
||
|
||
- **host** — one llama-server router endpoint (`base_url`, list of model ids
|
||
it can serve, per-model `parallel`, a speed `weight`).
|
||
- **route** — a client identity = the first path segment of the request URL
|
||
(`/opencode-a/v1/...` → route `opencode-a`). The client only ever knows a
|
||
base URL, so this needs no client support beyond "configurable base URL".
|
||
- **fingerprint** — a per-conversation key inside a route: SHA-256 of the
|
||
request's system prompt plus its first user message (first 4 KiB of each).
|
||
Stable across a conversation's turns, different between conversations. Empty
|
||
when the body has no user message (probes, title generation).
|
||
- **lease** — `(route, fingerprint, model) → host`, with `created`, `last_used`,
|
||
`state ∈ {active, pinned, draining}`. Pins are per route; a pinned route
|
||
pins all its fingerprints.
|
||
|
||
## 4. Request flow
|
||
|
||
```
|
||
client ── /{route}/v1/chat/completions ──▶ crossbar
|
||
1. route := first path segment; strip it
|
||
2. model := body.model (JSON peek; fall back to route default)
|
||
2b. fp := fingerprint(body) ("" if none)
|
||
3. lease := table[(route, fp, model)] → fallback table[(route, "", model)]
|
||
hit & host healthy & model resident → use it
|
||
miss | host unhealthy → choose(host) ; write lease
|
||
4. acquire one concurrency token for (host, model) [bounded queue]
|
||
5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
|
||
flush on every chunk; passive health on connect error / 5xx
|
||
6. release token; lease.last_used = now
|
||
```
|
||
|
||
`choose(host)`: among hosts that are healthy **and list the model as loaded**
|
||
(`/models`), take the one with the most free slots × `weight`; tie → lowest
|
||
current queue depth. If no host has it loaded, take a healthy host that *can*
|
||
serve it (config) and accept the load; log that decision.
|
||
|
||
Pass-through per route, all answered from the leased host (or from config if
|
||
no lease yet): `/v1/models`, `/health`, `/props`, `/v1/embeddings`,
|
||
`/v1/completions`. Everything else 404.
|
||
|
||
## 4a. Why two keys (Kyle's main use case: many OpenCode instances on one box)
|
||
|
||
Per-instance identity comes from the route; per-session identity comes from
|
||
the fingerprint. OpenCode's system prompt is per project and its first user
|
||
message is per session, so `(route, fp)` separates sessions inside one
|
||
instance without any client support. N sessions then spread across hosts by
|
||
free slots × weight *at start* and are held there; the per-(host, model) queue
|
||
(§6) is what stops N from oversubscribing any one host.
|
||
|
||
Client launcher (OpenCode config substitutes `{env:VAR}` in values — verified
|
||
in the installed 1.15.10 bundle: `/\{env:([^}]+)\}/g` → `process.env[V]`,
|
||
missing → empty string, so always set the variable; as a fallback the same
|
||
value can go through `options.headers["X-Crossbar-Route"]`, which crossbar
|
||
also accepts):
|
||
|
||
```jsonc
|
||
// ~/.config/opencode/opencode.json (one block for every project)
|
||
"provider": { "crossbar": { "npm": "@ai-sdk/openai-compatible",
|
||
"options": { "baseURL": "http://crossbar:7777/{env:CROSSBAR_ROUTE}/v1" },
|
||
"models": { "ornith-1.5-35b-a3b": {}, "laguna-s-2.1": {} } } }
|
||
```
|
||
```sh
|
||
# oc: one route per instance
|
||
CROSSBAR_ROUTE="$(basename "$PWD")-$$" exec opencode "$@"
|
||
```
|
||
|
||
## 4b. Context sizes: learned, not configured
|
||
|
||
The poller records `n_ctx`, `n_parallel` (→ per-slot context with unified KV)
|
||
from `/props` per host and model into `host_health`; `/props` is passed
|
||
through per route from the leased host so clients see the real number. Config
|
||
may override (`ctx = N` under a model) but normally does not. v2 uses it as a
|
||
guard: a body whose estimated prompt size (bytes/4 × 1.2) exceeds the leased
|
||
host's per-slot context is re-leased to a host where it fits, or answered
|
||
400 with a clear message instead of the upstream "Context size has been
|
||
exceeded".
|
||
|
||
## 5. Lease rules
|
||
|
||
- A lease is **sticky**. It moves only when: the host fails health, the route
|
||
is idle longer than `lease_idle` (default 30 min), or an operator pins/
|
||
releases it. A faster host coming back online does *not* move an active
|
||
lease — the re-prefill is the cost we are avoiding.
|
||
- **Pinned** leases never move automatically ("project A goes to titan right
|
||
now"). **Draining** hosts accept no new leases; existing ones finish.
|
||
- Table lives in memory and is written through to SQLite (§7a) on every
|
||
change; loaded at start so a crossbar restart does not reshuffle sessions.
|
||
|
||
## 6. Health
|
||
|
||
- Poller, every `poll_interval` (60 s): `GET /health` then `GET /models` per
|
||
host. Record loaded models. Only then `GET /slots?model=X` **for loaded
|
||
models only** — probing an unloaded model makes llama-server load it.
|
||
Per-model free-slot count from `/slots` (or, if `/slots` is disabled on a
|
||
host, assume `parallel` minus our own in-flight count).
|
||
- Passive: a connection error or 5xx on a proxied request marks the host
|
||
unhealthy immediately and re-leases the route on the client's retry.
|
||
Recovery only through the poller (two consecutive good polls).
|
||
- Concurrency: per `(host, model)` token bucket of size `parallel`. Requests
|
||
beyond it wait in a bounded FIFO (`queue_max`, default 8; 503 beyond that).
|
||
This is where the unified-KV overflow is prevented.
|
||
|
||
## 7. Admin API (tailnet only, same listener, prefix `/_crossbar`)
|
||
|
||
- `GET /_crossbar/hosts` — health, loaded models, free slots, in-flight.
|
||
- `GET /_crossbar/routes` — the lease table.
|
||
- `POST /_crossbar/routes/{route}` `{ "host": "titan", "pin": true }` — pin;
|
||
`{ "release": true }` — drop the lease (next request re-chooses).
|
||
- `POST /_crossbar/hosts/{host}` `{ "drain": true|false }`.
|
||
- `GET /_crossbar/usage?since=…&by=route|model|host` — accounting rollups
|
||
from §7a (JSON; `Accept: text/plain` gives a table).
|
||
- `GET /_crossbar/metrics` — Prometheus: requests, queue wait, lease moves,
|
||
host health, tokens/s from llama-server's `timings` when present. Scrape it
|
||
from the fleet Prometheus on orion.
|
||
|
||
## 7a. State store and accounting (SQLite)
|
||
|
||
One SQLite file (`crossbar.db`, WAL mode) holds both the durable state and the
|
||
accounting log. Driver: `modernc.org/sqlite` (pure Go, no cgo) so the arm64
|
||
static build stays a plain `go build`. Single writer goroutine fed by a
|
||
channel; readers use their own connection. Volume is a few rows per request,
|
||
so nothing here is performance-sensitive.
|
||
|
||
```sql
|
||
CREATE TABLE leases ( -- current table, one row per (route, model)
|
||
route TEXT, fp TEXT, model TEXT, host TEXT, state TEXT, -- active|pinned|draining
|
||
created INTEGER, last_used INTEGER, PRIMARY KEY (route, fp, model));
|
||
|
||
CREATE TABLE lease_events ( -- why sessions moved
|
||
ts INTEGER, route TEXT, model TEXT, from_host TEXT, to_host TEXT,
|
||
reason TEXT); -- new|unhealthy|idle|pin|release|drain
|
||
|
||
CREATE TABLE requests ( -- one row per proxied completion
|
||
id INTEGER PRIMARY KEY, route TEXT, fp TEXT, model TEXT, host TEXT,
|
||
started INTEGER, queued_ms INTEGER, ttfb_ms INTEGER, total_ms INTEGER,
|
||
status INTEGER, streamed INTEGER,
|
||
prompt_tokens INTEGER, cached_tokens INTEGER, completion_tokens INTEGER,
|
||
err TEXT);
|
||
|
||
CREATE TABLE host_health ( -- poller observations, for uptime accounting
|
||
ts INTEGER, host TEXT, healthy INTEGER, loaded_models TEXT);
|
||
```
|
||
|
||
Token and cache figures come from the upstream response when llama-server
|
||
provides them: `usage` on non-streaming replies, and the final SSE chunk's
|
||
`usage` / `timings` (`prompt_n`, `cache_n`, `predicted_n`, `predicted_ms`) on
|
||
streamed replies. To see that chunk the proxy tees the response body through
|
||
a small SSE line scanner; it never buffers or alters the stream. When the
|
||
fields are absent, wall-clock columns are still recorded.
|
||
|
||
What this answers: per route (session/project), per model, per host — number
|
||
of requests, busy seconds (`sum(total_ms)`), tokens in/out, cache-hit ratio
|
||
(`cached_tokens / prompt_tokens`, the direct measure of whether affinity is
|
||
working), queue wait, error rate, and per-host uptime. `/_crossbar/usage`
|
||
exposes the rollups; a nightly job prunes `requests` older than `retention`
|
||
(default 180 d) into a `requests_daily` rollup so the file stays small.
|
||
|
||
Body contents are never stored — only counts and timings.
|
||
|
||
## 8. Config (`crossbar.toml`)
|
||
|
||
TOML (`github.com/BurntSushi/toml`): no indentation semantics, no implicit
|
||
type coercion, and it is what Kyle's Rust projects already use. A NixOS module
|
||
can generate it with `pkgs.formats.toml`.
|
||
|
||
```toml
|
||
listen = "100.x.y.z:7777" # tailnet address only; never 0.0.0.0
|
||
db = "/var/lib/crossbar/crossbar.db"
|
||
poll_interval = "60s"
|
||
lease_idle = "30m"
|
||
queue_max = 8
|
||
retention = "180d"
|
||
|
||
[hosts.straylight]
|
||
base_url = "http://straylight.<tailnet>:11434"
|
||
weight = 1.0
|
||
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "ornith-1.5-9b-uncensored" = { parallel = 6 } }
|
||
|
||
[hosts.titan]
|
||
base_url = "http://titan.<tailnet>:8081"
|
||
weight = 2.0 # ~2x straylight decode
|
||
models = { "ornith-1.5-35b-a3b" = { parallel = 4 }, "laguna-s-2.1" = { parallel = 2 } }
|
||
|
||
[hosts.dixie]
|
||
base_url = "http://dixie.<tailnet>:11434"
|
||
weight = 0.8
|
||
models = { "ornith-1.5-9b-uncensored" = { parallel = 6 } }
|
||
|
||
# optional per-route defaults / pins
|
||
[routes.opencode-a]
|
||
default_model = "ornith-1.5-35b-a3b"
|
||
[routes.opencode-b]
|
||
default_model = "ornith-1.5-35b-a3b"
|
||
[routes.paper]
|
||
default_model = "qwen3.8-27b-uncensored"
|
||
pin = "titan"
|
||
[routes.hermes-straylight]
|
||
[routes.hermes-talos]
|
||
[routes.hermes-titan]
|
||
pin = "titan" # titan's models are the titan agent's first
|
||
```
|
||
|
||
Client side, no code changes:
|
||
|
||
- OpenCode: project-local `opencode.json` provider with
|
||
`options.baseURL: http://crossbar:7777/opencode-a/v1` (global + project
|
||
configs merge, so each project carries its own route).
|
||
- Hermes: `custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1`
|
||
(and the `delegation` / `auxiliary` blocks that point at a router today).
|
||
- tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry
|
||
the llama routers already have (`plain_http_to_sink` for that host only).
|
||
|
||
## 9. Security notes
|
||
|
||
- Bind to the tailnet address only. v0 relies on the tailnet for
|
||
authentication; every route is reachable by every tailnet peer, which is the
|
||
same exposure the llama-servers have today. v1 option: read Tailscale
|
||
identity (`tailscale whois` on the peer address) and restrict routes to
|
||
peers, so `hermes-talos` can only be used from talos.
|
||
- Admin API: same listener, same trust. Consider a separate `admin_listen` on
|
||
localhost if crossbar runs on a shared host.
|
||
- No secrets in config; llama-servers take no keys.
|
||
- Request bodies are proxied, never logged. Metrics carry counts and
|
||
timings only.
|
||
|
||
## 10. Milestones
|
||
|
||
- **v0 (a day):** TOML config, static routes → host mapping, reverse proxy with
|
||
streaming, `/health` + `/models` poller, passive health, `/_crossbar/hosts`.
|
||
Replaces the hand-maintained provider lists. No leases yet: each route has
|
||
a fixed host list in preference order; first healthy wins.
|
||
- **v1:** SQLite state + accounting (§7a), lease table keyed by
|
||
`(route, fingerprint, model)`, header route override, `choose()` by free slots × weight,
|
||
per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
|
||
- **v2:** wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤
|
||
45 s, else fall through); context-size guard (§4b); Tailscale-identity
|
||
route restriction.
|
||
|
||
## 11. Testing
|
||
|
||
- `httptest` fake llama-server: `/health`, `/models`, `/slots`, streaming
|
||
`/v1/chat/completions` with configurable latency and failure injection.
|
||
- Table tests for `choose()`, lease expiry, drain, passive-health re-lease,
|
||
fingerprint stability (same conversation → same key; different first user
|
||
message → different key; no user message → empty key).
|
||
- SSE tee scanner: fixture streams with and without a final `usage` chunk;
|
||
assert the client receives the bytes unchanged and the row is recorded.
|
||
- One integration script against the real fleet: one long conversation,
|
||
assert every turn hits the same host (llama-server `timings.cache_n` or
|
||
`prompt_n` small after the first turn).
|
||
|
||
## 12. Open questions
|
||
|
||
1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static
|
||
Go binary + systemd unit; arm64 cross-compile is free. Candidate:
|
||
hyperborea, in its own unit, outside the collector's disk budget.
|
||
2. `/slots` is disabled by default on recent llama-server builds; enable
|
||
`--slots` on the three routers or rely on our own in-flight counts.
|
||
3. Should `default_model` rewrite a client's `model` field? Proposal: no —
|
||
proxy what the client sent; only use the default when the body has none.
|
||
4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does
|
||
titan become a general pool member now that health is tracked? Kyle said
|
||
titan could join the pool (2026-09-25); the pin above keeps its own agent
|
||
first without excluding others.
|