v2.3 implementation plan: clients that manage their own slots
For the implementing model: do not work from this file. The owner gives you one task file at a time. This file is the index for the owner and the reviewer.
Goal: serve Boxmaker, a harness whose inferproxy talks plain HTTP/1.1 to one host:port and
rewrites nothing. It pins id_slot, polls GET /slots?model= while it waits, reads
GET /props?model= once, and sends POST /v1/chat/completions/control on a second connection
while its own stream is running. Checked on 2026-09-25 against crossbar at 4c64158, it failed on
six counts (thread i7jeubrtziru38s5gn8gmha44a): no route in its paths; /slots and /tokenize
not proxied; side calls leased separately from the stream; /control taking a limiter slot
behind its own stream; crossbar's queue hiding a waiting request from the server's /slots; and a
context refusal that is not llama-server's exceed_context_size_error.
- 01-control-plane — every route: a GET's model comes from
?model=;/slotsand/tokenizeare proxied; control calls (any GET/HEAD,POST /tokenize,POST /v1/chat/completions/control) follow the lease but skip the limiter, the context guard and the accounting row. Given:proxy/control_test.go; replacesproxy/proxy_test.go(v1: the/r/slots → 404row becomes/r/slots/0and/r/metrics). - 02-affinity-queue — route keys
affinity = "route"(one lease for the route) andqueue = false(count the request as load, never hold or refuse it);limiter.Track. Given:config/config_v23_test.go,limiter/track_test.go,proxy/affinity_test.go. - 03-route-listeners — route key
listen: a dedicated listener where every request is that route with an unprefixed path;Handler.ForRoute,identity.RouteMiddleware, one server per listener inmain. Given:config/listen_test.go,proxy/listener_test.go,identity/route_middleware_test.go; replacesexample.toml(v2.2: addsboxmaker-a) andtools/smoke.sh(v2: adds check 6, the dedicated listener). - 04-ctx-error-docs — the context refusal in llama-server's shape
{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":N,"n_ctx":M}}; README. Given: replacesproxy/ctxguard_test.go(v2).
Order matters: 02's affinity test uses /slots (01); 03's listener test uses route affinity
(02). Each task is green on its own given tests plus all earlier ones.
How this plan was made: acceptance tests first, no reference implementation; the given tests
compiled against a panic-only skeleton of the new names (Route.PerRoute, Route.Queues,
Route.Listen, Limiter.Track, Handler.ForRoute, identity.RouteMiddleware) on master
4c64158 and failed there for the intended reasons (404 on /slots//tokenize, 503 queue full
on control calls, 4 stray accounting rows, the old error body, requests held behind one slot,
unvalidated listen/affinity).
Facts about the live hosts (2026-09-25): all three routers run llama-server b10964; /slots
answers 200 on all three; POST /v1/chat/completions/control exists ({"success":false,"message":"no active completion for this id"} for an unknown id). In router mode GET /slots?model=X and
/props?model=X autoload X — a control call only ever reaches the leased host, which is where
the client's chat goes anyway, so this is the load the client asked for.
Global constraints
- Everything in
AGENTS.md. Branchv2.3frommaster. One task, one fresh OpenCode session, one commit. Given files are copied and never edited; earlier plans' given files stay protected, except the four this plan replaces (proxy/proxy_test.go,proxy/ctxguard_test.go,example.toml,tools/smoke.sh), whose v2.3 copies are then the protected ones.
Changes during the run
(none yet)