Context refusal in llama-server's exceed_context_size_error shape; README for v2.3
Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
This commit is contained in:
@@ -59,6 +59,9 @@ hosts = ["beta", "alpha"]
|
||||
| `hosts.<name>.models` | The models this host serves, with per-model parallel tuning. |
|
||||
| `routes.<name>.hosts` | Candidate hosts, tried in order until one is healthy; a conversation leases one of them. |
|
||||
| `routes.<name>.default_model` | Model used when a request omits one; must be served by a host in the route. |
|
||||
| `routes.<name>.affinity` | `"conversation"` (default, one lease per conversation) or `"route"` (one lease for the whole route); see "Clients that manage their own slots". |
|
||||
| `routes.<name>.queue` | `false` leaves queueing to the client's own llama-server slot; the default counts requests in crossbar's per-(host, model) queue. |
|
||||
| `routes.<name>.listen` | A host:port for the route's own listener, every request there is this route; see "Clients that manage their own slots". |
|
||||
| `identity` | `"off"` (default), `"tailscale"`, or `"header"`; see below. |
|
||||
| `hosts.<name>.wake` | A wake-on-LAN target (`mac`, `broadcast`, `wait`) so crossbar can rouse a sleeping host when nothing else can take a new lease. |
|
||||
| `routes.<name>.peers` | The tailnet nodes allowed to reach the route, with `identity = "tailscale"`; see below. |
|
||||
@@ -104,6 +107,46 @@ curl -H 'X-Crossbar-Route: opencode-a' \
|
||||
https://crossbar.<tailnet>:7777/v1/chat/completions
|
||||
```
|
||||
|
||||
## Clients that manage their own slots
|
||||
|
||||
Some clients connect to one crossbar address and manage a llama-server slot themselves: they pin
|
||||
`id_slot`, poll `/slots`, and steer a running completion through
|
||||
`/v1/chat/completions/control`. `inferproxy`, Boxmaker's router, is one. crossbar serves such a
|
||||
client from a route that has its own `listen` address and `affinity = "route"`, so the whole route
|
||||
lives on one host:
|
||||
|
||||
```toml
|
||||
# a client that manages its own llama-server slot (it pins id_slot, polls /slots, steers a
|
||||
# running completion through /v1/chat/completions/control) and cannot put a route in the path.
|
||||
# The route gets its own port; every request there is this route and the path goes upstream as is.
|
||||
[routes.boxmaker-a]
|
||||
hosts = ["beta", "alpha"]
|
||||
default_model = "ornith-1.5-35b-a3b"
|
||||
listen = "127.0.0.1:17801" # a tailnet address in production; never the main listen address
|
||||
affinity = "route" # one lease for the whole route, not one per conversation
|
||||
queue = false # counted as load but never held or refused: the server's own slot queue does that
|
||||
```
|
||||
|
||||
Every request to that address is this route, with its whole path passed upstream unchanged (there is
|
||||
no route segment to strip), so it runs through `Handler.ForRoute` rather than the usual
|
||||
`/{route}/` path. The address must split into a host and a numeric port, be unique across routes,
|
||||
not equal the top-level `listen`, and not be on a template route — crossbar refuses any of those at
|
||||
start-up.
|
||||
|
||||
A few things about how crossbar treats those requests:
|
||||
|
||||
- **Control calls take no slot.** A GET or HEAD on any allowed path, and a POST to exactly
|
||||
`/tokenize` or `/v1/chat/completions/control`, is a control call. It follows the route's single
|
||||
lease but takes no slot, skips the context guard, and writes no accounting row: it is sent beside
|
||||
its own stream, so it must never wait for or hold a slot. A chat completion on `/v1/chat/completions`
|
||||
is not a control call.
|
||||
- **`/slots` and `/tokenize` are proxied; `/slots/<id>` actions are not.** Only the bare `/slots`
|
||||
path is allowed, so an action on a specific slot id is not forwarded.
|
||||
- **A GET's model comes from its `?model=` query** (there is no body to read), which is how
|
||||
`/slots?model=shared` learns which model's slots to report.
|
||||
- **The admin API is not served on a route listener.** `/_crossbar/hosts` there, and any prefixed
|
||||
path such as `/boxmaker-a/v1/models`, are 404.
|
||||
|
||||
## Operate
|
||||
|
||||
The operator's API lives under `/_crossbar/`. Every call returns 200 with a small JSON body unless
|
||||
@@ -175,9 +218,14 @@ with the leased host's per-slot context for that model. A prompt that fits stays
|
||||
does not fit is moved to a healthy host on the route where it does fit (the lease moves with
|
||||
it, so the conversation stays there), and the response carries
|
||||
`X-Crossbar-Ctx: moved:<from>` + `>` + `<to>` — for example `moved:small>big`. When no host can
|
||||
fit it, the answer is `400 {"error":"prompt too large","estimate":<tokens>,"max":<largest per-slot
|
||||
context among hosts that have the model loaded>}`. Hosts whose context is unknown are never
|
||||
blocked by the guard.
|
||||
fit it, the answer is a `400` in llama-server's own overflow shape, so a client that handles the
|
||||
server's error handles crossbar's refusal too:
|
||||
|
||||
```json
|
||||
{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":<tokens>,"n_ctx":<largest per-slot context among hosts that have the model loaded>}}
|
||||
```
|
||||
|
||||
Hosts whose context is unknown are never blocked by the guard.
|
||||
|
||||
## Wake
|
||||
|
||||
@@ -204,4 +252,4 @@ tests only, so crossbar logs a warning when it starts in that mode.
|
||||
|
||||
## What v2 does not do
|
||||
|
||||
Request coalescing, `/slots` and TLS are out of scope for v2; see `PLAN.md`.
|
||||
Request coalescing and TLS are out of scope for v2; see `PLAN.md`.
|
||||
|
||||
Reference in New Issue
Block a user