README: context guard section (X-Crossbar-Ctx, 400 body), per-model context in the hosts view, router-mode note
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -110,12 +110,17 @@ The operator's API lives under `/_crossbar/`. Every call returns 200 with a smal
|
||||
stated otherwise.
|
||||
|
||||
`GET /_crossbar/hosts` reports every host's health, loaded models, live concurrency from the
|
||||
limiter and drain state:
|
||||
limiter, drain state and the context sizes the poller learned (`n_ctx`/`slots` from a single
|
||||
server's `/props`, `models` per loaded model from `/props?model=`; 0 or absent means unknown):
|
||||
|
||||
```json
|
||||
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false}}
|
||||
{"alpha":{"healthy":true,"loaded":["ornith-1.5-35b-a3b","small-9b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":7,"in_flight":0,"queued":0,"draining":false,"n_ctx":0,"slots":0,"models":{"ornith-1.5-35b-a3b":{"n_ctx":262144,"slots":4},"small-9b":{"n_ctx":32768,"slots":2}}},"beta":{"healthy":true,"loaded":["ornith-1.5-35b-a3b"],"last_ok":"2026-09-25T13:53:25Z","last_err":"","free_slots":2,"in_flight":0,"queued":0,"draining":false,"n_ctx":131072,"slots":2,"models":{}}}
|
||||
```
|
||||
|
||||
On a llama-server **router** only models whose `status.value` is `"loaded"` count as loaded, and
|
||||
crossbar asks `/props?model=X` only for those: asking about an unloaded model would make the
|
||||
router load it.
|
||||
|
||||
`GET /_crossbar/routes` reports each route's candidate hosts, default model, any pin and its live
|
||||
leases:
|
||||
|
||||
@@ -162,6 +167,18 @@ crossbar_host_healthy{host="alpha"} 1
|
||||
crossbar_host_healthy{host="beta"} 1
|
||||
```
|
||||
|
||||
## Context guard
|
||||
|
||||
With unified KV a host's usable context per request is its context size divided by its slots.
|
||||
crossbar estimates a chat request's size from its body (bytes/4 with a margin) and compares it
|
||||
with the leased host's per-slot context for that model. A prompt that fits stays put. One that
|
||||
does not fit is moved to a healthy host on the route where it does fit (the lease moves with
|
||||
it, so the conversation stays there), and the response carries
|
||||
`X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can
|
||||
fit it, the answer is `400 {"error":"prompt too large","estimate":<tokens>,"max":<largest per-slot
|
||||
context among hosts that have the model loaded>}`. Hosts whose context is unknown are never
|
||||
blocked by the guard.
|
||||
|
||||
## Wake
|
||||
|
||||
When a route has no healthy host left and at least one candidate lists a `wake` target, crossbar
|
||||
|
||||
Reference in New Issue
Block a user