# v2.3 implementation plan: clients that manage their own slots > **For the implementing model:** do not work from this file. The owner gives you one task file at > a time. This file is the index for the owner and the reviewer. **Goal:** serve Boxmaker, a harness whose `inferproxy` talks plain HTTP/1.1 to one host:port and rewrites nothing. It pins `id_slot`, polls `GET /slots?model=` while it waits, reads `GET /props?model=` once, and sends `POST /v1/chat/completions/control` on a second connection while its own stream is running. Checked on 2026-09-25 against crossbar at 4c64158, it failed on six counts (thread `i7jeubrtziru38s5gn8gmha44a`): no route in its paths; `/slots` and `/tokenize` not proxied; side calls leased separately from the stream; `/control` taking a limiter slot behind its own stream; crossbar's queue hiding a waiting request from the server's `/slots`; and a context refusal that is not llama-server's `exceed_context_size_error`. - **01-control-plane** — every route: a GET's model comes from `?model=`; `/slots` and `/tokenize` are proxied; control calls (any GET/HEAD, `POST /tokenize`, `POST /v1/chat/completions/control`) follow the lease but skip the limiter, the context guard and the accounting row. Given: `proxy/control_test.go`; replaces `proxy/proxy_test.go` (v1: the `/r/slots → 404` row becomes `/r/slots/0` and `/r/metrics`). - **02-affinity-queue** — route keys `affinity = "route"` (one lease for the route) and `queue = false` (count the request as load, never hold or refuse it); `limiter.Track`. Given: `config/config_v23_test.go`, `limiter/track_test.go`, `proxy/affinity_test.go`. - **03-route-listeners** — route key `listen`: a dedicated listener where every request is that route with an unprefixed path; `Handler.ForRoute`, `identity.RouteMiddleware`, one server per listener in `main`. Given: `config/listen_test.go`, `proxy/listener_test.go`, `identity/route_middleware_test.go`; replaces `example.toml` (v2.2: adds `boxmaker-a`) and `tools/smoke.sh` (v2: adds check 6, the dedicated listener). - **04-ctx-error-docs** — the context refusal in llama-server's shape `{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":N,"n_ctx":M}}`; README. Given: replaces `proxy/ctxguard_test.go` (v2). **Order matters:** 02's affinity test uses `/slots` (01); 03's listener test uses route affinity (02). Each task is green on its own given tests plus all earlier ones. **How this plan was made:** acceptance tests first, no reference implementation; the given tests compiled against a panic-only skeleton of the new names (`Route.PerRoute`, `Route.Queues`, `Route.Listen`, `Limiter.Track`, `Handler.ForRoute`, `identity.RouteMiddleware`) on master 4c64158 and failed there for the intended reasons (404 on `/slots`/`/tokenize`, 503 queue full on control calls, 4 stray accounting rows, the old error body, requests held behind one slot, unvalidated `listen`/`affinity`). **Facts about the live hosts (2026-09-25):** all three routers run llama-server b10964; `/slots` answers 200 on all three; `POST /v1/chat/completions/control` exists (`{"success":false,"message":"no active completion for this id"}` for an unknown id). In router mode `GET /slots?model=X` and `/props?model=X` **autoload X** — a control call only ever reaches the leased host, which is where the client's chat goes anyway, so this is the load the client asked for. ## Global constraints - Everything in `AGENTS.md`. Branch `v2.3` from `master`. One task, one fresh OpenCode session, one commit. Given files are copied and never edited; earlier plans' given files stay protected, except the four this plan replaces (`proxy/proxy_test.go`, `proxy/ctxguard_test.go`, `example.toml`, `tools/smoke.sh`), whose v2.3 copies are then the protected ones. ## Changes during the run - 2026-09-25, before task 02: straylight ran short of memory and Claude Code's reaper killed the driver after task 01 committed (`33fa61b`, first-gate); resumed at 02 an hour later. - Task 02: **owner test fault, model hack.** `TestQueueFalseNeitherHoldsNorRefuses` checked `InFlight == 0` right after the answers arrived, but the slot is released by a deferred call just after the answer is sent. Ornith "fixed" the race by releasing the slot at the first `Flush` — for every route, so a streaming request stopped counting against the limit at its first byte (the limiter no longer limited generation). No given test caught it. Fixed the test (waits for the release) and added `TestLoadIsHeldForTheWholeStream` (reads the first SSE chunk, asserts the slot is still held; fails on the hack, passes without it). Owner removed the `onFlush` hook and the `release` parameter from `forward`. The session then ended on a refused `/tmp` write while committing (refusal-ending #9); owner committed its staged work.