Files
crossbar/PLAN.md
T

9.0 KiB
Raw Blame History

crossbar — an affinity router for the fleet's llama-servers

Status: plan, 2026-09-25. Nothing built. Go. Owner: Kyle. Drafted by claude from the 2026-09-25 discussion (Inference Infrastructure).

1. Problem

Three hosts run llama-server in router mode with the same model ids (straylight, titan, soon dixie). Clients (OpenCode per project, Hermes agents, the paper writer) each hard-code one host. Consequences seen this week:

  • a host's prompt cache is per process; an agent session sits at 100–230k tokens with ~100 % cache hits. Any move between hosts re-prefills the whole context (≈ 7–8 min at straylight's ~500 t/s prompt processing);
  • two turns landing on one 35B at once overflowed its unified KV ("Context size has been exceeded"); nothing in the path knows the slot count;
  • titan sleeps; clients pointed at it fail instead of falling back;
  • rebuilding or restarting a host takes its clients down with it.

A crossbar switch connects each caller to a line and holds the connection. That is the design: affinity first, balancing only for new sessions.

2. Non-goals

  • No token counting, prompt rewriting, batching or splitting across hosts. The hosts are bandwidth-bound; there is nothing to gain cross-host.
  • No model management. Loading/unloading stays with each host's llama-server router (--models-max). crossbar never sends a request that would make a host load a model it does not have resident.
  • No auth in v0 (tailnet-only bind). See §9.

3. Terminology

  • host — one llama-server router endpoint (base_url, list of model ids it can serve, per-model parallel, a speed weight).
  • route — a client identity = the first path segment of the request URL (/opencode-a/v1/... → route opencode-a). The client only ever knows a base URL, so this needs no client support beyond "configurable base URL".
  • lease — (route, model) → host, with created, last_used, state ∈ {active, pinned, draining}.

4. Request flow

client  ── /{route}/v1/chat/completions ──▶ crossbar
        1. route := first path segment; strip it
        2. model := body.model (JSON peek; fall back to route default)
        3. lease := table[(route, model)]
             hit  & host healthy & model resident → use it
             miss | host unhealthy               → choose(host) ; write lease
        4. acquire one concurrency token for (host, model)   [bounded queue]
        5. reverse-proxy to host.base_url + "/v1/..." ; stream through,
           flush on every chunk; passive health on connect error / 5xx
        6. release token; lease.last_used = now

choose(host): among hosts that are healthy and list the model as loaded (/models), take the one with the most free slots × weight; tie → lowest current queue depth. If no host has it loaded, take a healthy host that can serve it (config) and accept the load; log that decision.

Pass-through per route, all answered from the leased host (or from config if no lease yet): /v1/models, /health, /props, /v1/embeddings, /v1/completions. Everything else 404.

5. Lease rules

  • A lease is sticky. It moves only when: the host fails health, the route is idle longer than lease_idle (default 30 min), or an operator pins/ releases it. A faster host coming back online does not move an active lease — the re-prefill is the cost we are avoiding.
  • Pinned leases never move automatically ("project A goes to titan right now"). Draining hosts accept no new leases; existing ones finish.
  • Table persisted to state.json after every change; loaded at start so a crossbar restart does not reshuffle sessions.

6. Health

  • Poller, every poll_interval (60 s): GET /health then GET /models per host. Record loaded models. Only then GET /slots?model=X for loaded models only — probing an unloaded model makes llama-server load it. Per-model free-slot count from /slots (or, if /slots is disabled on a host, assume parallel minus our own in-flight count).
  • Passive: a connection error or 5xx on a proxied request marks the host unhealthy immediately and re-leases the route on the client's retry. Recovery only through the poller (two consecutive good polls).
  • Concurrency: per (host, model) token bucket of size parallel. Requests beyond it wait in a bounded FIFO (queue_max, default 8; 503 beyond that). This is where the unified-KV overflow is prevented.

7. Admin API (tailnet only, same listener, prefix /_crossbar)

  • GET /_crossbar/hosts — health, loaded models, free slots, in-flight.
  • GET /_crossbar/routes — the lease table.
  • POST /_crossbar/routes/{route} { "host": "titan", "pin": true } — pin; { "release": true } — drop the lease (next request re-chooses).
  • POST /_crossbar/hosts/{host} { "drain": true|false }.
  • GET /_crossbar/metrics — Prometheus: requests, queue wait, lease moves, host health, tokens/s from llama-server's timings when present. Scrape it from the fleet Prometheus on orion.

8. Config (crossbar.yaml)

listen: "100.x.y.z:7777"        # tailnet address only; never 0.0.0.0
poll_interval: 60s
lease_idle: 30m
queue_max: 8
hosts:
  straylight:
    base_url: http://straylight.<tailnet>:11434
    weight: 1.0
    models: { ornith-1.5-35b-a3b: {parallel: 4}, ornith-1.5-9b-uncensored: {parallel: 6} }
  titan:
    base_url: http://titan.<tailnet>:8081
    weight: 2.0                 # ~2× straylight decode
    models: { ornith-1.5-35b-a3b: {parallel: 4}, laguna-s-2.1: {parallel: 2} }
  dixie:
    base_url: http://dixie.<tailnet>:11434
    weight: 0.8
    models: { ornith-1.5-9b-uncensored: {parallel: 6} }
routes:                          # optional defaults / pins
  opencode-a: { default_model: ornith-1.5-35b-a3b }
  opencode-b: { default_model: ornith-1.5-35b-a3b }
  paper:      { default_model: qwen3.8-27b-uncensored, pin: titan }
  hermes-straylight: {}
  hermes-talos: {}
  hermes-titan: { pin: titan }   # titan's models are the titan agent's first

Client side, no code changes:

  • OpenCode: project-local opencode.json provider with options.baseURL: http://crossbar:7777/opencode-a/v1 (global + project configs merge, so each project carries its own route).
  • Hermes: custom_providers[].base_url: http://crossbar:7777/hermes-<agent>/v1 (and the delegation / auxiliary blocks that point at a router today).
  • tirith: crossbar's plain-HTTP tailnet URL needs the same narrow trust entry the llama routers already have (plain_http_to_sink for that host only).

9. Security notes

  • Bind to the tailnet address only. v0 relies on the tailnet for authentication; every route is reachable by every tailnet peer, which is the same exposure the llama-servers have today. v1 option: read Tailscale identity (tailscale whois on the peer address) and restrict routes to peers, so hermes-talos can only be used from talos.
  • Admin API: same listener, same trust. Consider a separate admin_listen on localhost if crossbar runs on a shared host.
  • No secrets in config; llama-servers take no keys.
  • Request bodies are proxied, never logged. Metrics carry counts and timings only.

10. Milestones

  • v0 (a day): config, static routes → host mapping, reverse proxy with streaming, /health + /models poller, passive health, /_crossbar/hosts. Replaces the hand-maintained provider lists. No leases yet: each route has a fixed host list in preference order; first healthy wins.
  • v1: lease table with persistence, choose() by free slots × weight, per-(host, model) concurrency + bounded queue, pin/release/drain, metrics.
  • v2: wake-on-LAN for a sleeping titan when a new lease wants it (wait ≤ 45 s, else fall through); optional conversation fingerprint (system prompt + first user message) for per-conversation stickiness inside one route; Tailscale-identity route restriction.

11. Testing

  • httptest fake llama-server: /health, /models, /slots, streaming /v1/chat/completions with configurable latency and failure injection.
  • Table tests for choose(), lease expiry, drain, passive-health re-lease.
  • One integration script against the real fleet: one long conversation, assert every turn hits the same host (llama-server timings.cache_n or prompt_n small after the first turn).

12. Open questions

  1. Where it runs — Kyle: "a Raspberry Pi, it doesn't matter where". Static Go binary + systemd unit; arm64 cross-compile is free. Candidate: hyperborea, in its own unit, outside the collector's disk budget.
  2. /slots is disabled by default on recent llama-server builds; enable --slots on the three routers or rely on our own in-flight counts.
  3. Should default_model rewrite a client's model field? Proposal: no — proxy what the client sent; only use the default when the body has none.
  4. Does the titan agent's reservation (2026-09-21) stand as a pin, or does titan become a general pool member now that health is tracked? Kyle said titan could join the pool (2026-09-25); the pin above keeps its own agent first without excluding others.