Files
boxmaker/docs/plans/M2a/README.md
T

5.1 KiB

M2a implementation plan: the inference path

For the implementing model: do not work from this file. The owner gives you one task file at a time (01-… to 13-…). This file is the index for the owner and the reviewer.

Goal: loopd can hold a correct, robust conversation with llama-server through a Unix socket: requests built from typed input, streams reassembled exactly, every kind of silence and failure ending in a defined way, and a self-test that refuses a server it does not recognise.

Architecture: inferproxy forwards bytes between infer.sock and the router. In loopd, a hand-written HTTP/1.1 client and SSE reader sit under a llama client that owns the waits, the liveness limit, the thinking cap, the per-slot gate and retry. Nothing in the client starts a thread; timers are socket read timeouts. Behaviour is pinned by tests that run against a scripted fake server replaying responses recorded from straylight.

Tech stack: Rust stable (edition 2024, rust-version = "1.95"), serde, serde_json, toml, emsha 1.0.4 (new), std only for sockets and threads.

Spec: docs/specs/2026-09-17-m2a-inference-path.md. Measurements: docs/inference-contract.md.

Global constraints

  • Everything in AGENTS.md, including "Lessons from earlier reviews".
  • No new dependency except emsha. No HTTP, SSE, async or randomness crate.
  • Library code never panics on what a peer sends. Everything read from a socket is bounded.
  • Our own formats reject unknown fields. The server's JSON does not: it sends fields we do not use and newer builds add more, so structs that parse the server's responses must not use deny_unknown_fields. Each task says which kind it is dealing with.
  • Branch m2a. One task, one fresh OpenCode session, one commit. Run cargo fmt --all before the gate. Review happens once, after task 13.

Tasks

# File Delivers Tests that define it
01 01-proto-sha256.md proto::sha256, Sha256, HashError over emsha proto/tests/hash.rs
02 02-inferproxy.md The forwarder, its two limits, its command line inferproxy/tests/bucket.rs, forward.rs
03 03-loopd-config.md loopd::config loopd/tests/config.rs
04 04-loopd-http.md loopd::http; brings in the fake server and all recordings loopd/tests/http.rs
05 05-loopd-sse.md loopd::sse loopd/tests/sse.rs
06 06-llama-request.md loopd::llama types and request::build_body loopd/tests/request.rs
07 07-llama-assemble.md llama::assemble::Assembler loopd/tests/assemble.rs
08 08-llama-info.md props, slots, tokenize, cache_outcome loopd/tests/info.rs
09 09-llama-chat.md Client::chat: waits, liveness, errors loopd/tests/chat.rs
10 10-llama-cap.md The thinking cap loopd/tests/cap.rs
11 11-llama-gate-retry.md SlotGate, chat_with_retry loopd/tests/retry.rs
12 12-selftest.md loopd::selftest, loopd selftest --config loopd/tests/selftest.rs
13 13-verify-device.md make verify-device against straylight loopd/tests/device.rs

files/ holds everything the tasks copy into place: tests, the fake server (loopd/tests/support/mod.rs), recordings (fixtures/http/*.http), expected results derived from the recordings by a separate script (fixtures/expected/), and the new Makefile.

The plan was checked the same way as M1: a private reference implementation passes every test, the gate passes after every task in order, the timing tests were run repeatedly under CPU load, and the reference passes make verify-device on straylight.

For the owner: running a task

In ~/src/boxmaker, start a fresh OpenCode session with Laguna S 2.1 and send:

Read docs/plans/M2a/01-proto-sha256.md and do exactly that task.

Then the next file in a new session. Or let tools/run-plan.sh docs/plans/M2a do that: it runs each task in a fresh opencode run session with Laguna, and stops at the first task that does not end with a commit, a clean tree and a done row. If a session ends with a stopped row in docs/implementer-log.md, do not start the next task. Task 13 talks to straylight: Ornith must be loadable and slot 0 should not be in heavy use while it runs.

For the reviewer: after task 13

  1. git log --oneline master..m2a: thirteen commits with the Implemented-By trailer.
  2. Copied files are unchanged: for f in $(cd docs/plans/M2a/files && find . -type f); do cmp "docs/plans/M2a/files/$f" "$f"; done
  3. git diff master..m2a --stat -- docs/design.md docs/specs docs/plans AGENTS.md CLAUDE.md deny.toml is empty.
  4. make gate, make audit, make verify-device.
  5. Read every source file against its task and the spec. Probe from outside with inputs the tests do not contain, especially: malformed HTTP, a server that misbehaves mid-stream, several threads on the gate, and limits at their boundaries.
  6. Run the timing tests repeatedly under CPU load.
  7. Write findings under "Reviews" in docs/implementer-log.md, and turn them into rows in docs/implementer-lessons.md, filling in "Seen again" for the M1 tips.