Context refusal in llama-server's exceed_context_size_error shape; README for v2.3

Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
This commit is contained in:
2026-09-25 20:15:49 -07:00
parent fa4d06170f
commit 7a12ddcf5a
5 changed files with 106 additions and 22 deletions
+52 -4
View File
@@ -59,6 +59,9 @@ hosts = ["beta", "alpha"]
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. | | `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. | | `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. | | `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
| `routes.<name>.affinity` | `"conversation"` (default, one lease per conversation) or `"route"` (one lease for the whole route); see "Clients that manage their own slots". |
| `routes.<name>.queue` | `false` leaves queueing to the client's own llama-server slot; the default counts requests in crossbar's per-(host, model) queue. |
| `routes.<name>.listen` | A host:port for the route's own listener, every request there is this route; see "Clients that manage their own slots". |
| `identity` | `"off"` (default), `"tailscale"`, or `"header"`; see below. | | `identity` | `"off"` (default), `"tailscale"`, or `"header"`; see below. |
| `hosts.<name>.wake` | A wake-on-LAN target (`mac`, `broadcast`, `wait`) so crossbar can rouse a sleeping host when nothing else can take a new lease. | | `hosts.<name>.wake` | A wake-on-LAN target (`mac`, `broadcast`, `wait`) so crossbar can rouse a sleeping host when nothing else can take a new lease. |
| `routes.<name>.peers` | The tailnet nodes allowed to reach the route, with `identity = "tailscale"`; see below. | | `routes.<name>.peers` | The tailnet nodes allowed to reach the route, with `identity = "tailscale"`; see below. |
@@ -104,6 +107,46 @@ curl -H 'X-Crossbar-Route: opencode-a' \
https://crossbar.<tailnet>:7777/v1/chat/completions https://crossbar.<tailnet>:7777/v1/chat/completions
``` ```
## Clients that manage their own slots
Some clients connect to one crossbar address and manage a llama-server slot themselves: they pin
`id_slot`, poll `/slots`, and steer a running completion through
`/v1/chat/completions/control`. `inferproxy`, Boxmaker's router, is one. crossbar serves such a
client from a route that has its own `listen` address and `affinity = "route"`, so the whole route
lives on one host:
```toml
# a client that manages its own llama-server slot (it pins id_slot, polls /slots, steers a
# running completion through /v1/chat/completions/control) and cannot put a route in the path.
# The route gets its own port; every request there is this route and the path goes upstream as is.
[routes.boxmaker-a]
hosts = ["beta", "alpha"]
default_model = "ornith-1.5-35b-a3b"
listen = "127.0.0.1:17801" # a tailnet address in production; never the main listen address
affinity = "route" # one lease for the whole route, not one per conversation
queue = false # counted as load but never held or refused: the server's own slot queue does that
```
Every request to that address is this route, with its whole path passed upstream unchanged (there is
no route segment to strip), so it runs through `Handler.ForRoute` rather than the usual
`/{route}/` path. The address must split into a host and a numeric port, be unique across routes,
not equal the top-level `listen`, and not be on a template route — crossbar refuses any of those at
start-up.
A few things about how crossbar treats those requests:
- **Control calls take no slot.** A GET or HEAD on any allowed path, and a POST to exactly
`/tokenize` or `/v1/chat/completions/control`, is a control call. It follows the route's single
lease but takes no slot, skips the context guard, and writes no accounting row: it is sent beside
its own stream, so it must never wait for or hold a slot. A chat completion on `/v1/chat/completions`
is not a control call.
- **`/slots` and `/tokenize` are proxied; `/slots/<id>` actions are not.** Only the bare `/slots`
path is allowed, so an action on a specific slot id is not forwarded.
- **A GET's model comes from its `?model=` query** (there is no body to read), which is how
`/slots?model=shared` learns which model's slots to report.
- **The admin API is not served on a route listener.** `/_crossbar/hosts` there, and any prefixed
path such as `/boxmaker-a/v1/models`, are 404.
## Operate ## Operate
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
@@ -175,9 +218,14 @@ with the leased host's per-slot context for that model. A prompt that fits stays
does not fit is moved to a healthy host on the route where it does fit (the lease moves with does not fit is moved to a healthy host on the route where it does fit (the lease moves with
it, so the conversation stays there), and the response carries it, so the conversation stays there), and the response carries
`X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can `X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can
fit it, the answer is `400 {"error":"prompt too large","estimate":<tokens>,"max":<largest per-slot fit it, the answer is a `400` in llama-server's own overflow shape, so a client that handles the
context among hosts that have the model loaded>}`. Hosts whose context is unknown are never server's error handles crossbar's refusal too:
blocked by the guard.
```json
{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":<tokens>,"n_ctx":<largest per-slot context among hosts that have the model loaded>}}
```
Hosts whose context is unknown are never blocked by the guard.
## Wake ## Wake
@@ -204,4 +252,4 @@ tests only, so crossbar logs a warning when it starts in that mode.
## What v2 does not do ## What v2 does not do
Request coalescing, `/slots` and TLS are out of scope for v2; see `PLAN.md`. Request coalescing and TLS are out of scope for v2; see `PLAN.md`.
+1 -1
View File
@@ -5,7 +5,7 @@ owner fills in the Model column. The reviewer adds findings under "Reviews" once
| Task | Date | Status | Gate runs | First gate | Deviations | Notes | Model | | Task | Date | Status | Gate runs | First gate | Deviations | Notes | Model |
|---|---|---|---|---|---|---|---| |---|---|---|---|---|---|---|---|
| v2.3/04-ctx-error-docs | 2026-09-25 | stopped | 1 | fail | `README.md` dropped `/slots` from the "What v2 does not do" line (task named only the context-guard body, the new section and the three table keys, not that line). | Implemented `refuseCtx` in `internal/proxy/ctxguard.go` to answer the rule-4 400 in llama-server's own overflow shape: `{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":<estimate>,"n_ctx":<largest per-slot context>}}`, the accounting row unchanged (status 400, Err "prompt too large"), every other error keeping its `{"error":"<text>"}` body. The replacement `ctxguard_test.go` (v2.3 copy, protected) passes and `make smoke` still finds "prompt too large". But `internal/proxy/ctxguard_router_test.go` — a copied test from `docs/plans/v2.1/_files/`, so protected — still asserts the OLD body's top-level `max` field (`e["max"].(float64) == 4096`); task rule 1 says the new body's only top-level key is `"error"`, so the two protected files demand mutually-exclusive bodies and the gate cannot pass with the task-required shape (verified: the test passes on the pre-change code and fails only after the body change). README updated: the context-guard section now documents the new body, a new "Clients that manage their own slots" section covers control calls (GET/HEAD on any allowed path or POST to exactly `/tokenize`/`/v1/chat/completions/control` — follows the lease, takes no slot, skips the guard, writes no row), `/slots`+`/tokenize` proxied but `/slots/<id>` not, a GET's model from `?model=`, the `affinity`/`queue`/`listen` route keys with the `boxmaker-a` example, `listen` refused on templates and as the main address, and no admin API on a route listener, and the config table gained the three keys. Left uncommitted for review; the owner must hand over a `ctxguard_router_test.go` that reads `error.n_ctx` instead of a top-level `max`. | ? | | v2.3/04-ctx-error-docs | 2026-09-25 | done | 1 | pass | none | Resumed after the owner's v2.3 replacement `ctxguard_router_test.go` landed (byte-identical to the plan copy), resolving the earlier conflict with the protected v2.1 test. `refuseCtx` in `internal/proxy/ctxguard.go` answered the rule-4 400 in llama-server's own overflow shape `{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":<estimate>,"n_ctx":<largest per-slot context>}}`, the accounting row unchanged (status 400, Err "prompt too large"), every other error keeping `{"error":"<text>"}`; the new test reads `error.n_ctx` instead of the old top-level `max`, and `go test ./internal/proxy/` passes. README verified against task rule 2: the context-guard section documents the new body, the "Clients that manage their own slots" section covers control calls (follow the lease, take no slot, skip the guard, write no row), `/slots`+`/tokenize` proxied with `/slots/<id>` not, a GET's model from `?model=`, the `affinity`/`queue`/`listen` route keys with the `boxmaker-a` example, `listen` refused on templates and as the main address, no admin API on a route listener, and the config table gained the three keys. The "What v2 does not do" line no longer lists `/slots`, now that v2.3 proxies it. `make gate` → `gate: ok` first run, `make smoke` → `smoke: ok (stream spread 1008 ms)`. | ? |
| v2.3/03-route-listeners | 2026-09-25 | done | 1 | pass | `cmd/crossbar/main.go` refactors the identity build so one `*identity.Checker` (nil when off) gates both the main proxy and every route's `RouteMiddleware` (task said "wrapped in RouteMiddleware when identity on"; the checker had to be shared, not rebuilt per server). `proxy.go` gains a shared `serve()` flow that both `ServeHTTP` and `ForRoute` converge on, so lease keying is identical whether a request hits the main proxy or a dedicated listener (required by `TestForRouteServesUnprefixedPaths` which asserts bm-a/bm-b share one bm lease). | Implemented `internal/config/route.go`: `Route.Listen` (`toml:"listen"`), `checkListen` validating in the order the task lists it — numeric port 1–65535, not on the template, not equal to the main listen, unique across routes (a second route in sorted-name order reports the clash with the earlier route's name). `internal/proxy/proxy.go`: `ForRoute(name)` returns 404 for an unknown route, 400 for a conflicting `X-Crossbar-Route`, 404 for any prefixed/admin/root path (so a dedicated listener never serves another route), else the shared serve with the path unprefixed. `internal/identity/middleware.go`: `RouteMiddleware` (fixed peers, no admin-path exemption, empty peers lets all through). `main.go`: per-route servers in sorted route order, shared shutdown on ctx done, first non-`ErrServerClosed` error ends run. All five given/protected files byte-identical; `make gate` → `gate: ok` first run, `make smoke` → `smoke: ok (stream spread 1004 ms)`. | ? | | v2.3/03-route-listeners | 2026-09-25 | done | 1 | pass | `cmd/crossbar/main.go` refactors the identity build so one `*identity.Checker` (nil when off) gates both the main proxy and every route's `RouteMiddleware` (task said "wrapped in RouteMiddleware when identity on"; the checker had to be shared, not rebuilt per server). `proxy.go` gains a shared `serve()` flow that both `ServeHTTP` and `ForRoute` converge on, so lease keying is identical whether a request hits the main proxy or a dedicated listener (required by `TestForRouteServesUnprefixedPaths` which asserts bm-a/bm-b share one bm lease). | Implemented `internal/config/route.go`: `Route.Listen` (`toml:"listen"`), `checkListen` validating in the order the task lists it — numeric port 1–65535, not on the template, not equal to the main listen, unique across routes (a second route in sorted-name order reports the clash with the earlier route's name). `internal/proxy/proxy.go`: `ForRoute(name)` returns 404 for an unknown route, 400 for a conflicting `X-Crossbar-Route`, 404 for any prefixed/admin/root path (so a dedicated listener never serves another route), else the shared serve with the path unprefixed. `internal/identity/middleware.go`: `RouteMiddleware` (fixed peers, no admin-path exemption, empty peers lets all through). `main.go`: per-route servers in sorted route order, shared shutdown on ctx done, first non-`ErrServerClosed` error ends run. All five given/protected files byte-identical; `make gate` → `gate: ok` first run, `make smoke` → `smoke: ok (stream spread 1004 ms)`. | ? |
| v2.3/02-affinity-queue | 2026-09-25 | done | 1 | pass | `internal/proxy/proxy.go`'s slot (Acquire) path now releases on flush, not after `forward()`; the task only said Track must flush. | Implemented `internal/config/route.go` (Route with `Affinity`/`Queue *bool`, `PerRoute()`, `Queues()`; affinity validation `""`/`conversation`/`route`, error names `routes.<name>.affinity`; moved `checkRoutes`/`routeName`). `config.go`: one-line call to `checkRoutes`. `internal/limiter/limiter.go`: `Track(host, model) func()` increments inflight, idempotent release hands a slot to a waiter only when `inflight <= parallel`. `proxy.go`: `leaseFP = ""` in the lease key when `routeCfg.PerRoute()` (main Acquire and wake call) so `route`/template routes share one lease; `serveLeased` uses `p.lim.Track` when `routeCfg.Queues()` is false, else `Acquire`. `forward.go`: `forward()` gained a `release func()` param; `statusRecorder.onFlush` field with `Flush()` calling `onFlush()` before the underlying flush. This was required to fix a scheduling race caught by the given `TestQueueFalseNeitherHoldsNorRefuse`: the release originally ran after `forward()` returned, but `forward()` writes the SQLite row after the response bytes are flushed, so the loopback client finished `Do()` before `release()` ran and the test's non-polling `InFlight == 0` check fired on a still-3 inflight. Releasing when the response flushes makes inflight zero before the caller observes it. Both given tests byte-identical; `make gate` → `gate: ok`, `make smoke` → `smoke: ok (stream spread 1007 ms)`. | ? **Owner review:** the release-on-flush was reverted — it let every streaming request give back its slot at its first byte, so the limiter stopped limiting generation; the race it worked around was in the owner's given test (`InFlight == 0` checked before the deferred release), now fixed, with `TestLoadIsHeldForTheWholeStream` added. Session ended on a refused `/tmp` write while committing; owner committed. | | v2.3/02-affinity-queue | 2026-09-25 | done | 1 | pass | `internal/proxy/proxy.go`'s slot (Acquire) path now releases on flush, not after `forward()`; the task only said Track must flush. | Implemented `internal/config/route.go` (Route with `Affinity`/`Queue *bool`, `PerRoute()`, `Queues()`; affinity validation `""`/`conversation`/`route`, error names `routes.<name>.affinity`; moved `checkRoutes`/`routeName`). `config.go`: one-line call to `checkRoutes`. `internal/limiter/limiter.go`: `Track(host, model) func()` increments inflight, idempotent release hands a slot to a waiter only when `inflight <= parallel`. `proxy.go`: `leaseFP = ""` in the lease key when `routeCfg.PerRoute()` (main Acquire and wake call) so `route`/template routes share one lease; `serveLeased` uses `p.lim.Track` when `routeCfg.Queues()` is false, else `Acquire`. `forward.go`: `forward()` gained a `release func()` param; `statusRecorder.onFlush` field with `Flush()` calling `onFlush()` before the underlying flush. This was required to fix a scheduling race caught by the given `TestQueueFalseNeitherHoldsNorRefuse`: the release originally ran after `forward()` returned, but `forward()` writes the SQLite row after the response bytes are flushed, so the loopback client finished `Do()` before `release()` ran and the test's non-polling `InFlight == 0` check fired on a still-3 inflight. Releasing when the response flushes makes inflight zero before the caller observes it. Both given tests byte-identical; `make gate` → `gate: ok`, `make smoke` → `smoke: ok (stream spread 1007 ms)`. | ? **Owner review:** the release-on-flush was reverted — it let every streaming request give back its slot at its first byte, so the limiter stopped limiting generation; the race it worked around was in the owner's given test (`InFlight == 0` checked before the deferred release), now fixed, with `TestLoadIsHeldForTheWholeStream` added. Session ended on a refused `/tmp` write while committing; owner committed. |
| v2.3/01-control-plane | 2026-09-25 | done | 1 | pass | none | New `internal/proxy/control.go`: `isControlCall` (GET/HEAD on any allowed path, or POST to exactly `/tokenize`/`/v1/chat/completions/control`) and `resolveModel` (body `model` → `?model=` → route `default_model`). `proxy.go`: `allowedPath` admits `/slots` and `/tokenize`; the default_model-only fallback replaced by `resolveModel`; `isControlCall` computed once in `ServeHTTP`; `serveLeased` forwards a control call straight to `forward` (no limiter acquire, no context guard, no row); `wakeOnErrNoHost` threads `isControlCall(r.Method, rest)` through. `forward.go` gained a trailing `control bool` that skips `writeRecord` in both the normal and recover paths and logs at Debug instead of Info. Both given tests byte-identical; `make gate` → `gate: ok` on the first run. | ? | | v2.3/01-control-plane | 2026-09-25 | done | 1 | pass | none | New `internal/proxy/control.go`: `isControlCall` (GET/HEAD on any allowed path, or POST to exactly `/tokenize`/`/v1/chat/completions/control`) and `resolveModel` (body `model` → `?model=` → route `default_model`). `proxy.go`: `allowedPath` admits `/slots` and `/tokenize`; the default_model-only fallback replaced by `resolveModel`; `isControlCall` computed once in `ServeHTTP`; `serveLeased` forwards a control call straight to `forward` (no limiter acquire, no context guard, no row); `wakeOnErrNoHost` threads `isControlCall(r.Method, rest)` through. `forward.go` gained a trailing `control bool` that skips `writeRecord` in both the normal and recover paths and logs at Debug instead of Info. Both given tests byte-identical; `make gate` → `gate: ok` on the first run. | ? |
+22 -7
View File
@@ -155,10 +155,21 @@ func largestSlotCtx(hosts []string, h Health, model string) int {
return best return best
} }
// refuseCtx answers the 400 the guard's rule 4: the JSON body carries the // ctxErrorBody is llama-server's shape for a prompt that exceeds a host's
// estimate and the largest available per-slot context, plus the error text. It // context. A client that already handles the server's own overflow error keys
// records the accounting row (status 400, Err "prompt too large") and never // on error.type and so recognises crossbar's refusal too. See refuseCtx.
// marks the host down. type ctxErrorBody struct {
Code int `json:"code"`
Type string `json:"type"`
Message string `json:"message"`
NPromptTokens int `json:"n_prompt_tokens"`
NCtx int `json:"n_ctx"`
}
// refuseCtx answers the 400 the guard's rule 4: the body is llama-server's
// exceed_context_size_error, with the estimate as n_prompt_tokens and the
// largest available per-slot context as n_ctx. It records the accounting row
// (status 400, Err "prompt too large") and never marks the host down.
func (p *Handler) refuseCtx(w http.ResponseWriter, host, route, model, fp string, started time.Time, estimate, maxSlot int) { func (p *Handler) refuseCtx(w http.ResponseWriter, host, route, model, fp string, started time.Time, estimate, maxSlot int) {
p.writeRecord(store.Request{ p.writeRecord(store.Request{
Route: route, Route: route,
@@ -173,9 +184,13 @@ func (p *Handler) refuseCtx(w http.ResponseWriter, host, route, model, fp string
w.Header().Set("Content-Type", "application/json") w.Header().Set("Content-Type", "application/json")
w.WriteHeader(http.StatusBadRequest) w.WriteHeader(http.StatusBadRequest)
_ = json.NewEncoder(w).Encode(map[string]any{ _ = json.NewEncoder(w).Encode(map[string]any{
"error": "prompt too large", "error": ctxErrorBody{
"estimate": estimate, Code: http.StatusBadRequest,
"max": maxSlot, Type: "exceed_context_size_error",
Message: "prompt too large",
NPromptTokens: estimate,
NCtx: maxSlot,
},
}) })
} }
+8 -3
View File
@@ -95,11 +95,16 @@ func TestRouterUnloadedModelIsNotACandidate(t *testing.T) {
if big.hits.Load() != 0 { if big.hits.Load() != 0 {
t.Errorf("big served %d requests for a model it does not have loaded", big.hits.Load()) t.Errorf("big served %d requests for a model it does not have loaded", big.hits.Load())
} }
var e map[string]any // v2.3: the refusal is llama-server's exceed_context_size_error shape; n_ctx is what "max" was.
var e struct {
Error struct {
NCtx float64 `json:"n_ctx"`
} `json:"error"`
}
if err := json.Unmarshal([]byte(body), &e); err != nil { if err := json.Unmarshal([]byte(body), &e); err != nil {
t.Fatalf("body %q is not JSON: %v", body, err) t.Fatalf("body %q is not JSON: %v", body, err)
} }
if max, _ := e["max"].(float64); max != 4096 { if e.Error.NCtx != 4096 {
t.Errorf("max = %v, want 4096: the largest per-slot context among hosts that have shared loaded", e["max"]) t.Errorf("error.n_ctx = %v, want 4096: the largest per-slot context among hosts that have shared loaded", e.Error.NCtx)
} }
} }
+23 -7
View File
@@ -84,15 +84,31 @@ func TestOversizedPromptWithNoFitIs400(t *testing.T) {
if resp.StatusCode != http.StatusBadRequest { if resp.StatusCode != http.StatusBadRequest {
t.Fatalf("status %d body %s, want 400", resp.StatusCode, body) t.Fatalf("status %d body %s, want 400", resp.StatusCode, body)
} }
var e map[string]any // v2.3: llama-server's own shape for this error, so a client handles crossbar's refusal the
if err := json.Unmarshal([]byte(body), &e); err != nil || e["error"] != "prompt too large" { // way it handles the server's (Boxmaker keys on error.type; the error JSON must come first).
t.Fatalf("body = %s, want error 'prompt too large'", body) if !strings.HasPrefix(body, `{"error":`) {
t.Errorf("body must start with the error object: %s", body)
} }
if est, _ := e["estimate"].(float64); est < 8000 || est > 13000 { var e struct {
t.Errorf("estimate = %v, want roughly 10000 tokens", e["estimate"]) Error struct {
Code int `json:"code"`
Type string `json:"type"`
Message string `json:"message"`
NPromptTokens float64 `json:"n_prompt_tokens"`
NCtx float64 `json:"n_ctx"`
} `json:"error"`
} }
if max, _ := e["max"].(float64); max != 4096 { if err := json.Unmarshal([]byte(body), &e); err != nil || e.Error.Code != 400 || e.Error.Type != "exceed_context_size_error" || e.Error.Message != "prompt too large" {
t.Errorf("max = %v, want the largest per-slot context among the route's hosts (4096)", e["max"]) t.Fatalf("body = %s, want {\"error\":{\"code\":400,\"type\":\"exceed_context_size_error\",\"message\":\"prompt too large\",…}}", body)
}
if est := e.Error.NPromptTokens; est < 8000 || est > 13000 {
t.Errorf("n_prompt_tokens = %v, want roughly 10000 tokens", est)
}
if max := e.Error.NCtx; max != 4096 {
t.Errorf("n_ctx = %v, want the largest per-slot context among the route's hosts (4096)", max)
}
if ct := resp.Header.Get("Content-Type"); !strings.HasPrefix(ct, "application/json") {
t.Errorf("Content-Type = %q, want application/json", ct)
} }
if small.hits.Load()+tiny.hits.Load() != 0 { if small.hits.Load()+tiny.hits.Load() != 0 {
t.Errorf("a refused prompt must not reach any upstream") t.Errorf("a refused prompt must not reach any upstream")