Files
crossbar/docs/plans/v2.3/README.md
T

5.5 KiB

v2.3 implementation plan: clients that manage their own slots

For the implementing model: do not work from this file. The owner gives you one task file at a time. This file is the index for the owner and the reviewer.

Goal: serve Boxmaker, a harness whose inferproxy talks plain HTTP/1.1 to one host:port and rewrites nothing. It pins id_slot, polls GET /slots?model= while it waits, reads GET /props?model= once, and sends POST /v1/chat/completions/control on a second connection while its own stream is running. Checked on 2026-09-25 against crossbar at 4c64158, it failed on six counts (thread i7jeubrtziru38s5gn8gmha44a): no route in its paths; /slots and /tokenize not proxied; side calls leased separately from the stream; /control taking a limiter slot behind its own stream; crossbar's queue hiding a waiting request from the server's /slots; and a context refusal that is not llama-server's exceed_context_size_error.

  • 01-control-plane — every route: a GET's model comes from ?model=; /slots and /tokenize are proxied; control calls (any GET/HEAD, POST /tokenize, POST /v1/chat/completions/control) follow the lease but skip the limiter, the context guard and the accounting row. Given: proxy/control_test.go; replaces proxy/proxy_test.go (v1: the /r/slots → 404 row becomes /r/slots/0 and /r/metrics).
  • 02-affinity-queue — route keys affinity = "route" (one lease for the route) and queue = false (count the request as load, never hold or refuse it); limiter.Track. Given: config/config_v23_test.go, limiter/track_test.go, proxy/affinity_test.go.
  • 03-route-listeners — route key listen: a dedicated listener where every request is that route with an unprefixed path; Handler.ForRoute, identity.RouteMiddleware, one server per listener in main. Given: config/listen_test.go, proxy/listener_test.go, identity/route_middleware_test.go; replaces example.toml (v2.2: adds boxmaker-a) and tools/smoke.sh (v2: adds check 6, the dedicated listener).
  • 04-ctx-error-docs — the context refusal in llama-server's shape {"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":N,"n_ctx":M}}; README. Given: replaces proxy/ctxguard_test.go (v2) and proxy/ctxguard_router_test.go (v2.1).

Order matters: 02's affinity test uses /slots (01); 03's listener test uses route affinity (02). Each task is green on its own given tests plus all earlier ones.

How this plan was made: acceptance tests first, no reference implementation; the given tests compiled against a panic-only skeleton of the new names (Route.PerRoute, Route.Queues, Route.Listen, Limiter.Track, Handler.ForRoute, identity.RouteMiddleware) on master 4c64158 and failed there for the intended reasons (404 on /slots//tokenize, 503 queue full on control calls, 4 stray accounting rows, the old error body, requests held behind one slot, unvalidated listen/affinity).

Facts about the live hosts (2026-09-25): all three routers run llama-server b10964; /slots answers 200 on all three; POST /v1/chat/completions/control exists ({"success":false,"message":"no active completion for this id"} for an unknown id). In router mode GET /slots?model=X and /props?model=X autoload X — a control call only ever reaches the leased host, which is where the client's chat goes anyway, so this is the load the client asked for.

Global constraints

  • Everything in AGENTS.md. Branch v2.3 from master. One task, one fresh OpenCode session, one commit. Given files are copied and never edited; earlier plans' given files stay protected, except the five this plan replaces (proxy/proxy_test.go, proxy/ctxguard_test.go, proxy/ctxguard_router_test.go, example.toml, tools/smoke.sh), whose v2.3 copies are then the protected ones.

Changes during the run

  • 2026-09-25, before task 02: straylight ran short of memory and Claude Code's reaper killed the driver after task 01 committed (33fa61b, first-gate); resumed at 02 an hour later.
  • Task 02: owner test fault, model hack. TestQueueFalseNeitherHoldsNorRefuses checked InFlight == 0 right after the answers arrived, but the slot is released by a deferred call just after the answer is sent. Ornith "fixed" the race by releasing the slot at the first Flush — for every route, so a streaming request stopped counting against the limit at its first byte (the limiter no longer limited generation). No given test caught it. Fixed the test (waits for the release) and added TestLoadIsHeldForTheWholeStream (reads the first SSE chunk, asserts the slot is still held; fails on the hack, passes without it). Owner removed the onFlush hook and the release parameter from forward. The session then ended on a refused /tmp write while committing (refusal-ending #9); owner committed its staged work.
  • Task 03: first session emitted a stray </tool_call> after reading files and ended with no change (model); restarted unchanged, done in 16 min (fa1c398).
  • Task 04: owner fault, correct stop. The v2.1 given ctxguard_router_test.go also asserts the refusal body (e["max"]); I grepped only for the "prompt too large" string when writing the replacement list. Ornith implemented the new shape, saw the two protected tests demand incompatible bodies, committed only its stopped row, and reported — exactly the AGENTS.md rule. Replacement ctxguard_router_test.go (reads error.n_ctx) added.