Files
crossbar/docs/plans/v2/README.md
T

70 lines
4.2 KiB
Markdown

# v2 implementation plan: learned context, the context guard, wake-on-LAN, identity
> **For the implementing model:** do not work from this file. The owner gives you one task file at
> a time (`01-…` to `05-…`). This file is the index for the owner and the reviewer.
**Goal:** `PLAN.md` §4b and §10 v2. The poller learns each host's context size from `/props`; a
prompt that cannot fit the leased host's per-slot context moves to one where it fits or is
refused with a clear 400; a route whose hosts are all down can wake a sleeping host by
wake-on-LAN and wait for it; a route can be restricted to named tailnet peers.
**Architecture:** `health.Status` gains `NCtx`/`Slots` (task 01); the proxy gains the guard
(task 02); two new small packages, `wake` (magic packets + a waiter, task 03) and `identity`
(whois resolver, checker, middleware, task 04); config gains `identity`, `[hosts.x.wake]`,
`routes.x.peers`; `main` wires the waker and the middleware (task 05).
**How this plan was made:** acceptance tests first, from `PLAN.md`; no reference implementation.
Every given test compiled against a panic-only skeleton of the names in the tasks (`go vet`
clean). The given tests were walked against the task rules and against the other given files
(helpers, line limits, `main.go` call sites) before handover — the v1 findings list is the
reason.
**Tech stack:** as v1; no new module. `identity`'s production resolver shells out to
`tailscale whois --json`, which exists on every fleet host.
## Global constraints
- Everything in `AGENTS.md`. Branch `v2`. One task, one fresh OpenCode session, one commit.
- Bodies never logged. Type assertions two-valued. Files under 400 lines.
- Given files are copied and never edited; some **replace** earlier ones (the task says so).
## Tasks
| # | File | Delivers | Tests that define it |
|---|---|---|---|
| 01 | `01-props.md` | `Status.NCtx`, `Status.Slots`, `PerSlotCtx()`; `/props` in the poll; hosts view shows them | `health/props_test.go` |
| 02 | `02-ctxguard.md` | prompt-size estimate; move or 400; `X-Crossbar-Ctx` | `proxy/ctxguard_test.go` |
| 03 | `03-wake.md` | `internal/wake`: magic packet, `Send`, `Waker` | `wake/wake_test.go` |
| 04 | `04-identity.md` | `internal/identity`: whois parse, checker, header mode, middleware; config `identity`/`peers`/`wake` | `identity/*_test.go`, `config/config_v2_test.go` |
| 05 | `05-wiring-smoke.md` | proxy wakes on no-host; `main` wires waker + middleware; given fakeupstream/smoke/example; README | `make smoke` |
## For the owner
`tools/run-plan.sh docs/plans/v2` from a clean checkout on `master`.
## For the reviewer: after task 05
1. Five task commits with the trailer; given files byte-identical; protected files untouched.
2. `make gate`, `make smoke`.
3. Probe: a `/props` that returns 200 with a huge body (bounded read); a MAC with an unusual
separator in config; `identity = "tailscale"` on a host where `tailscale` is not on PATH
(must log and refuse the gated routes, never allow); two routes, one gated one open, from
the same peer; a wake target whose broadcast address is unroutable (503 within `wait`, no
hang); the guard with a body of exactly `MaxBody`.
4. Findings under "Reviews" in `docs/implementer-log.md`, by fault.
## Changes during the run
- 2026-09-25, task 01: the new `/props` poll lands on the v1 fake upstream's `/` catch-all, which
counts hits, so two v1 proxy tests with exact hit counts failed. Ornith implemented the task
correctly, did not touch the protected file, and stopped with a `stopped` row — exactly the
procedure. Owner's fault (T19 once more: a new task changed what an earlier given file
measures, and the pre-handover walk missed it). `helpers_test.go` is now a v2 given file that
answers `/props` without counting it; resumed.
- 2026-09-25, task 02: my `TestStickyLeaseSurvivesGrowthUntilItDoesNotFit` "grew" the conversation
by enlarging the *first user message*, which by the fingerprint spec makes it a different
conversation — so the test demanded `reused` for a new key. Ornith diagnosed it exactly ("turn
2's fp differs from turn 1's, yet the test expects reuse") and the session ended on a
malformed tool call. Test fault (mine): later turns are now appended after the first user
message. Resumed from the working tree.