Files
crossbar/docs/plans/v2.3/README.md
T

3.8 KiB

v2.3 implementation plan: clients that manage their own slots

For the implementing model: do not work from this file. The owner gives you one task file at a time. This file is the index for the owner and the reviewer.

Goal: serve Boxmaker, a harness whose inferproxy talks plain HTTP/1.1 to one host:port and rewrites nothing. It pins id_slot, polls GET /slots?model= while it waits, reads GET /props?model= once, and sends POST /v1/chat/completions/control on a second connection while its own stream is running. Checked on 2026-09-25 against crossbar at 4c64158, it failed on six counts (thread i7jeubrtziru38s5gn8gmha44a): no route in its paths; /slots and /tokenize not proxied; side calls leased separately from the stream; /control taking a limiter slot behind its own stream; crossbar's queue hiding a waiting request from the server's /slots; and a context refusal that is not llama-server's exceed_context_size_error.

  • 01-control-plane — every route: a GET's model comes from ?model=; /slots and /tokenize are proxied; control calls (any GET/HEAD, POST /tokenize, POST /v1/chat/completions/control) follow the lease but skip the limiter, the context guard and the accounting row. Given: proxy/control_test.go; replaces proxy/proxy_test.go (v1: the /r/slots → 404 row becomes /r/slots/0 and /r/metrics).
  • 02-affinity-queue — route keys affinity = "route" (one lease for the route) and queue = false (count the request as load, never hold or refuse it); limiter.Track. Given: config/config_v23_test.go, limiter/track_test.go, proxy/affinity_test.go.
  • 03-route-listeners — route key listen: a dedicated listener where every request is that route with an unprefixed path; Handler.ForRoute, identity.RouteMiddleware, one server per listener in main. Given: config/listen_test.go, proxy/listener_test.go, identity/route_middleware_test.go; replaces example.toml (v2.2: adds boxmaker-a) and tools/smoke.sh (v2: adds check 6, the dedicated listener).
  • 04-ctx-error-docs — the context refusal in llama-server's shape {"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":N,"n_ctx":M}}; README. Given: replaces proxy/ctxguard_test.go (v2).

Order matters: 02's affinity test uses /slots (01); 03's listener test uses route affinity (02). Each task is green on its own given tests plus all earlier ones.

How this plan was made: acceptance tests first, no reference implementation; the given tests compiled against a panic-only skeleton of the new names (Route.PerRoute, Route.Queues, Route.Listen, Limiter.Track, Handler.ForRoute, identity.RouteMiddleware) on master 4c64158 and failed there for the intended reasons (404 on /slots//tokenize, 503 queue full on control calls, 4 stray accounting rows, the old error body, requests held behind one slot, unvalidated listen/affinity).

Facts about the live hosts (2026-09-25): all three routers run llama-server b10964; /slots answers 200 on all three; POST /v1/chat/completions/control exists ({"success":false,"message":"no active completion for this id"} for an unknown id). In router mode GET /slots?model=X and /props?model=X autoload X — a control call only ever reaches the leased host, which is where the client's chat goes anyway, so this is the load the client asked for.

Global constraints

  • Everything in AGENTS.md. Branch v2.3 from master. One task, one fresh OpenCode session, one commit. Given files are copied and never edited; earlier plans' given files stay protected, except the four this plan replaces (proxy/proxy_test.go, proxy/ctxguard_test.go, example.toml, tools/smoke.sh), whose v2.3 copies are then the protected ones.

Changes during the run

(none yet)