Files
crossbar/docs/plans/v2

v2 implementation plan: learned context, the context guard, wake-on-LAN, identity

For the implementing model: do not work from this file. The owner gives you one task file at a time (01-… to 05-…). This file is the index for the owner and the reviewer.

Goal: PLAN.md §4b and §10 v2. The poller learns each host's context size from /props; a prompt that cannot fit the leased host's per-slot context moves to one where it fits or is refused with a clear 400; a route whose hosts are all down can wake a sleeping host by wake-on-LAN and wait for it; a route can be restricted to named tailnet peers.

Architecture: health.Status gains NCtx/Slots (task 01); the proxy gains the guard (task 02); two new small packages, wake (magic packets + a waiter, task 03) and identity (whois resolver, checker, middleware, task 04); config gains identity, [hosts.x.wake], routes.x.peers; main wires the waker and the middleware (task 05).

How this plan was made: acceptance tests first, from PLAN.md; no reference implementation. Every given test compiled against a panic-only skeleton of the names in the tasks (go vet clean). The given tests were walked against the task rules and against the other given files (helpers, line limits, main.go call sites) before handover — the v1 findings list is the reason.

Tech stack: as v1; no new module. identity's production resolver shells out to tailscale whois --json, which exists on every fleet host.

Global constraints

  • Everything in AGENTS.md. Branch v2. One task, one fresh OpenCode session, one commit.
  • Bodies never logged. Type assertions two-valued. Files under 400 lines.
  • Given files are copied and never edited; some replace earlier ones (the task says so).

Tasks

# File Delivers Tests that define it
01 01-props.md Status.NCtx, Status.Slots, PerSlotCtx(); /props in the poll; hosts view shows them health/props_test.go
02 02-ctxguard.md prompt-size estimate; move or 400; X-Crossbar-Ctx proxy/ctxguard_test.go
03 03-wake.md internal/wake: magic packet, Send, Waker wake/wake_test.go
04 04-identity.md internal/identity: whois parse, checker, header mode, middleware; config identity/peers/wake identity/*_test.go, config/config_v2_test.go
05 05-wiring-smoke.md proxy wakes on no-host; main wires waker + middleware; given fakeupstream/smoke/example; README make smoke

For the owner

tools/run-plan.sh docs/plans/v2 from a clean checkout on master.

For the reviewer: after task 05

  1. Five task commits with the trailer; given files byte-identical; protected files untouched.
  2. make gate, make smoke.
  3. Probe: a /props that returns 200 with a huge body (bounded read); a MAC with an unusual separator in config; identity = "tailscale" on a host where tailscale is not on PATH (must log and refuse the gated routes, never allow); two routes, one gated one open, from the same peer; a wake target whose broadcast address is unroutable (503 within wait, no hang); the guard with a body of exactly MaxBody.
  4. Findings under "Reviews" in docs/implementer-log.md, by fault.

Changes during the run

  • 2026-09-25, task 01: the new /props poll lands on the v1 fake upstream's / catch-all, which counts hits, so two v1 proxy tests with exact hit counts failed. Ornith implemented the task correctly, did not touch the protected file, and stopped with a stopped row — exactly the procedure. Owner's fault (T19 once more: a new task changed what an earlier given file measures, and the pre-handover walk missed it). helpers_test.go is now a v2 given file that answers /props without counting it; resumed.
  • 2026-09-25, task 02: my TestStickyLeaseSurvivesGrowthUntilItDoesNotFit "grew" the conversation by enlarging the first user message, which by the fingerprint spec makes it a different conversation — so the test demanded reused for a new key. Ornith diagnosed it exactly ("turn 2's fp differs from turn 1's, yet the test expects reuse") and the session ended on a malformed tool call. Test fault (mine): later turns are now appended after the first user message. Resumed from the working tree.
  • 2026-09-25, learned from titan's router (llama-server b10964) while v2 ran: in router mode a plain GET /props answers n_ctx: 0 (role: router), and GET /props?model=X autoloads X when models_autoload is on — the same trap as /slots?model=X. Task 01's poller therefore learns nothing on a real router and the guard stays inert there. Follow-up for v2.1: query /props?model=X only for models /v1/models lists as loaded, never for others. Not a defect in what the tasks asked for; a gap in what the owner knew when writing them.
  • 2026-09-25, task 05, first session: ended after 12 seconds. It misspelled the repository path (/home/kyle/src/crossar/Makefile), the sandbox refused the out-of-repository read, and it ended its turn — the sixth refusal-ending tonight, this one triggered by its own typo. Model fault; no change to the task. Restarted.
  • 2026-09-25, task 05, second session (30 min in, wiring written, smoke check 3 failing): the given tools/smoke.sh passed a 300 KB prompt as one curl -d argument, which Linux caps at 128 KiB per argv element, so check 3 could never pass. Test fault (mine): the body now goes through a file (-d @file). Ornith diagnosed it correctly. Resumed from the working tree with the corrected script.
  • 2026-09-25, task 05, third session: TestParallelAndQueue (v1 given limiter_test.go) failed once under full-suite -race load with after releases: inflight 1 queued 0. Test fault (mine): the third acquirer sent its result before its deferred release ran, so the final count check could observe one slot still held. The given file now releases before reporting. Ornith found it and measured the flake rate rather than editing the protected file.
  • 2026-09-25, task 05, third session: also saw TestQueueFullIs503 fail with Requests:3 Errors:2. Two causes. (1) Test fault (mine): arrival order rested on 30 ms sleeps; the given test now waits on the limiter's in-flight and queued counts. (2) A real v1 defect, verified by the owner with a diagnostic build (5 of 8 runs): after a forward completes, forward.go checks r.Context().Err() and, when the client has already closed its connection, records a served 200 as a 499 "client cancelled" error. Cancellation must be what the reverse proxy itself observed, never a post-hoc context check. Scheduled as v2.1 task 01; not fixed in task 05, which is wiring only. The session then ended on a refused read of /proc/loadavg — the seventh refusal-ending. Model fault.