README: context guard section (X-Crossbar-Ctx, 400 body), per-model context in the hosts view, router-mode note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
2026-09-25 11:50:38 -07:00
co-authored by Claude Fable 5.1
parent 11c8e053f7
commit 4f03cb2c52
+19 -2
View File
@@ -110,12 +110,17 @@ The operator's API lives under `/_crossbar/`. Every call returns 200 with a smal
stated otherwise. stated otherwise.
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the `GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
limiter and drain state: limiter, drain state and the context sizes the poller learned (`n_ctx`/`slots` from a single
server's `/props`, `models` per loaded model from `/props?model=`; 0 or absent means unknown):
```json ```json
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false}} {"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}}
``` ```
On a llama-server **router** only models whose `status.value` is `"loaded"` count as loaded, and
crossbar asks `/props?model=X` only for those: asking about an unloaded model would make the
router load it.
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live `GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
leases: leases:
@@ -162,6 +167,18 @@ crossbar_host_healthy{host="alpha"} 1
crossbar_host_healthy{host="beta"} 1 crossbar_host_healthy{host="beta"} 1
``` ```
## Context guard
With unified KV a host's usable context per request is its context size divided by its slots.
crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it
with the leased host's per-slot context for that model. A prompt that fits stays put. One that
does not fit is moved to a healthy host on the route where it does fit (the lease moves with
it, so the conversation stays there), and the response carries
`X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can
fit it, the answer is `400 {"error":"prompt too large","estimate":<tokens>,"max":<largest per-slot
context among hosts that have the model loaded>}`. Hosts whose context is unknown are never
blocked by the guard.
## Wake ## Wake
When a route has no healthy host left and at least one candidate lists a `wake` target, crossbar When a route has no healthy host left and at least one candidate lists a `wake` target, crossbar