Files
crossbar/docs/plans/v2.3/README.md
T

58 lines
3.8 KiB
Markdown

# v2.3 implementation plan: clients that manage their own slots
> **For the implementing model:** do not work from this file. The owner gives you one task file at
> a time. This file is the index for the owner and the reviewer.
**Goal:** serve Boxmaker, a harness whose `inferproxy` talks plain HTTP/1.1 to one host:port and
rewrites nothing. It pins `id_slot`, polls `GET /slots?model=` while it waits, reads
`GET /props?model=` once, and sends `POST /v1/chat/completions/control` on a second connection
while its own stream is running. Checked on 2026-09-25 against crossbar at 4c64158, it failed on
six counts (thread `i7jeubrtziru38s5gn8gmha44a`): no route in its paths; `/slots` and `/tokenize`
not proxied; side calls leased separately from the stream; `/control` taking a limiter slot
behind its own stream; crossbar's queue hiding a waiting request from the server's `/slots`; and a
context refusal that is not llama-server's `exceed_context_size_error`.
- **01-control-plane** — every route: a GET's model comes from `?model=`; `/slots` and
`/tokenize` are proxied; control calls (any GET/HEAD, `POST /tokenize`,
`POST /v1/chat/completions/control`) follow the lease but skip the limiter, the context guard
and the accounting row. Given: `proxy/control_test.go`; replaces `proxy/proxy_test.go` (v1:
the `/r/slots → 404` row becomes `/r/slots/0` and `/r/metrics`).
- **02-affinity-queue** — route keys `affinity = "route"` (one lease for the route) and
`queue = false` (count the request as load, never hold or refuse it); `limiter.Track`.
Given: `config/config_v23_test.go`, `limiter/track_test.go`, `proxy/affinity_test.go`.
- **03-route-listeners** — route key `listen`: a dedicated listener where every request is that
route with an unprefixed path; `Handler.ForRoute`, `identity.RouteMiddleware`, one server per
listener in `main`. Given: `config/listen_test.go`, `proxy/listener_test.go`,
`identity/route_middleware_test.go`; replaces `example.toml` (v2.2: adds `boxmaker-a`) and
`tools/smoke.sh` (v2: adds check 6, the dedicated listener).
- **04-ctx-error-docs** — the context refusal in llama-server's shape
`{"error":{"code":400,"type":"exceed_context_size_error","message":"prompt too large","n_prompt_tokens":N,"n_ctx":M}}`;
README. Given: replaces `proxy/ctxguard_test.go` (v2).
**Order matters:** 02's affinity test uses `/slots` (01); 03's listener test uses route affinity
(02). Each task is green on its own given tests plus all earlier ones.
**How this plan was made:** acceptance tests first, no reference implementation; the given tests
compiled against a panic-only skeleton of the new names (`Route.PerRoute`, `Route.Queues`,
`Route.Listen`, `Limiter.Track`, `Handler.ForRoute`, `identity.RouteMiddleware`) on master
4c64158 and failed there for the intended reasons (404 on `/slots`/`/tokenize`, 503 queue full
on control calls, 4 stray accounting rows, the old error body, requests held behind one slot,
unvalidated `listen`/`affinity`).
**Facts about the live hosts (2026-09-25):** all three routers run llama-server b10964; `/slots`
answers 200 on all three; `POST /v1/chat/completions/control` exists (`{"success":false,"message":"no
active completion for this id"}` for an unknown id). In router mode `GET /slots?model=X` and
`/props?model=X` **autoload X** — a control call only ever reaches the leased host, which is where
the client's chat goes anyway, so this is the load the client asked for.
## Global constraints
- Everything in `AGENTS.md`. Branch `v2.3` from `master`. One task, one fresh OpenCode session,
one commit. Given files are copied and never edited; earlier plans' given files stay protected,
except the four this plan replaces (`proxy/proxy_test.go`, `proxy/ctxguard_test.go`,
`example.toml`, `tools/smoke.sh`), whose v2.3 copies are then the protected ones.
## Changes during the run
(none yet)