# M2a implementation plan: the inference path > **For the implementing model:** do not work from this file. The owner gives you one task file at > a time (`01-…` to `13-…`). This file is the index for the owner and the reviewer. **Goal:** `loopd` can hold a correct, robust conversation with `llama-server` through a Unix socket: requests built from typed input, streams reassembled exactly, every kind of silence and failure ending in a defined way, and a self-test that refuses a server it does not recognise. **Architecture:** `inferproxy` forwards bytes between `infer.sock` and the router. In `loopd`, a hand-written HTTP/1.1 client and SSE reader sit under a llama client that owns the waits, the liveness limit, the thinking cap, the per-slot gate and retry. Nothing in the client starts a thread; timers are socket read timeouts. Behaviour is pinned by tests that run against a scripted fake server replaying responses recorded from straylight. **Tech stack:** Rust stable (edition 2024, `rust-version = "1.95"`), `serde`, `serde_json`, `toml`, `emsha` 1.0.4 (new), std only for sockets and threads. **Spec:** `docs/specs/2026-09-17-m2a-inference-path.md`. Measurements: `docs/inference-contract.md`. ## Global constraints - Everything in `AGENTS.md`, including "Lessons from earlier reviews". - No new dependency except `emsha`. No HTTP, SSE, async or randomness crate. - Library code never panics on what a peer sends. Everything read from a socket is bounded. - **Our own formats reject unknown fields. The server's JSON does not:** it sends fields we do not use and newer builds add more, so structs that parse the server's responses must not use `deny_unknown_fields`. Each task says which kind it is dealing with. - Branch `m2a`. One task, one fresh OpenCode session, one commit. Run `cargo fmt --all` before the gate. Review happens once, after task 13. ## Tasks | # | File | Delivers | Tests that define it | |---|---|---|---| | 01 | `01-proto-sha256.md` | `proto::sha256`, `Sha256`, `HashError` over `emsha` | `proto/tests/hash.rs` | | 02 | `02-inferproxy.md` | The forwarder, its two limits, its command line | `inferproxy/tests/bucket.rs`, `forward.rs` | | 03 | `03-loopd-config.md` | `loopd::config` | `loopd/tests/config.rs` | | 04 | `04-loopd-http.md` | `loopd::http`; brings in the fake server and all recordings | `loopd/tests/http.rs` | | 05 | `05-loopd-sse.md` | `loopd::sse` | `loopd/tests/sse.rs` | | 06 | `06-llama-request.md` | `loopd::llama` types and `request::build_body` | `loopd/tests/request.rs` | | 07 | `07-llama-assemble.md` | `llama::assemble::Assembler` | `loopd/tests/assemble.rs` | | 08 | `08-llama-info.md` | `props`, `slots`, `tokenize`, `cache_outcome` | `loopd/tests/info.rs` | | 09 | `09-llama-chat.md` | `Client::chat`: waits, liveness, errors | `loopd/tests/chat.rs` | | 10 | `10-llama-cap.md` | The thinking cap | `loopd/tests/cap.rs` | | 11 | `11-llama-gate-retry.md` | `SlotGate`, `chat_with_retry` | `loopd/tests/retry.rs` | | 12 | `12-selftest.md` | `loopd::selftest`, `loopd selftest --config` | `loopd/tests/selftest.rs` | | 13 | `13-verify-device.md` | `make verify-device` against straylight | `loopd/tests/device.rs` | `files/` holds everything the tasks copy into place: tests, the fake server (`loopd/tests/support/mod.rs`), recordings (`fixtures/http/*.http`), expected results derived from the recordings by a separate script (`fixtures/expected/`), and the new `Makefile`. The plan was checked the same way as M1: a private reference implementation passes every test, the gate passes after every task in order, the timing tests were run repeatedly under CPU load, and the reference passes `make verify-device` on straylight. ## For the owner: running a task In `~/src/boxmaker`, start a fresh OpenCode session with Laguna S 2.1 and send: > Read `docs/plans/M2a/01-proto-sha256.md` and do exactly that task. Then the next file in a new session. Or let `tools/run-plan.sh docs/plans/M2a` do that: it runs each task in a fresh `opencode run` session with Laguna, and stops at the first task that does not end with a commit, a clean tree and a `done` row. If a session ends with a `stopped` row in `docs/implementer-log.md`, do not start the next task. Task 13 talks to straylight: Ornith must be loadable and slot 0 should not be in heavy use while it runs. ## For the reviewer: after task 13 1. `git log --oneline master..m2a`: thirteen commits with the `Implemented-By` trailer. 2. Copied files are unchanged: `for f in $(cd docs/plans/M2a/files && find . -type f); do cmp "docs/plans/M2a/files/$f" "$f"; done` 3. `git diff master..m2a --stat -- docs/design.md docs/specs docs/plans AGENTS.md CLAUDE.md deny.toml` is empty. 4. `make gate`, `make audit`, `make verify-device`. 5. Read every source file against its task and the spec. Probe from outside with inputs the tests do not contain, especially: malformed HTTP, a server that misbehaves mid-stream, several threads on the gate, and limits at their boundaries. 6. Run the timing tests repeatedly under CPU load. 7. Write findings under "Reviews" in `docs/implementer-log.md`, and turn them into rows in `docs/implementer-lessons.md`, filling in "Seen again" for the M1 tips.