M2a implementation plan: the inference path
For the implementing model: do not work from this file. The owner gives you one task file at a time (
01-…to13-…). This file is the index for the owner and the reviewer.
Goal: loopd can hold a correct, robust conversation with llama-server through a Unix
socket: requests built from typed input, streams reassembled exactly, every kind of silence and
failure ending in a defined way, and a self-test that refuses a server it does not recognise.
Architecture: inferproxy forwards bytes between infer.sock and the router. In loopd, a
hand-written HTTP/1.1 client and SSE reader sit under a llama client that owns the waits, the
liveness limit, the thinking cap, the per-slot gate and retry. Nothing in the client starts a
thread; timers are socket read timeouts. Behaviour is pinned by tests that run against a scripted
fake server replaying responses recorded from straylight.
Tech stack: Rust stable (edition 2024, rust-version = "1.95"), serde, serde_json, toml,
emsha 1.0.4 (new), std only for sockets and threads.
Spec: docs/specs/2026-09-17-m2a-inference-path.md. Measurements:
docs/inference-contract.md.
Global constraints
- Everything in
AGENTS.md, including "Lessons from earlier reviews". - No new dependency except
emsha. No HTTP, SSE, async or randomness crate. - Library code never panics on what a peer sends. Everything read from a socket is bounded.
- Our own formats reject unknown fields. The server's JSON does not: it sends fields we do not
use and newer builds add more, so structs that parse the server's responses must not use
deny_unknown_fields. Each task says which kind it is dealing with. - Branch
m2a. One task, one fresh OpenCode session, one commit. Runcargo fmt --allbefore the gate. Review happens once, after task 13.
Tasks
| # | File | Delivers | Tests that define it |
|---|---|---|---|
| 01 | 01-proto-sha256.md |
proto::sha256, Sha256, HashError over emsha |
proto/tests/hash.rs |
| 02 | 02-inferproxy.md |
The forwarder, its two limits, its command line | inferproxy/tests/bucket.rs, forward.rs |
| 03 | 03-loopd-config.md |
loopd::config |
loopd/tests/config.rs |
| 04 | 04-loopd-http.md |
loopd::http; brings in the fake server and all recordings |
loopd/tests/http.rs |
| 05 | 05-loopd-sse.md |
loopd::sse |
loopd/tests/sse.rs |
| 06 | 06-llama-request.md |
loopd::llama types and request::build_body |
loopd/tests/request.rs |
| 07 | 07-llama-assemble.md |
llama::assemble::Assembler |
loopd/tests/assemble.rs |
| 08 | 08-llama-info.md |
props, slots, tokenize, cache_outcome |
loopd/tests/info.rs |
| 09 | 09-llama-chat.md |
Client::chat: waits, liveness, errors |
loopd/tests/chat.rs |
| 10 | 10-llama-cap.md |
The thinking cap | loopd/tests/cap.rs |
| 11 | 11-llama-gate-retry.md |
SlotGate, chat_with_retry |
loopd/tests/retry.rs |
| 12 | 12-selftest.md |
loopd::selftest, loopd selftest --config |
loopd/tests/selftest.rs |
| 13 | 13-verify-device.md |
make verify-device against straylight |
loopd/tests/device.rs |
files/ holds everything the tasks copy into place: tests, the fake server
(loopd/tests/support/mod.rs), recordings (fixtures/http/*.http), expected results derived from
the recordings by a separate script (fixtures/expected/), and the new Makefile.
The plan was checked the same way as M1: a private reference implementation passes every test, the
gate passes after every task in order, the timing tests were run repeatedly under CPU load, and the
reference passes make verify-device on straylight.
For the owner: running a task
In ~/src/boxmaker, start a fresh OpenCode session with Laguna S 2.1 and send:
Read
docs/plans/M2a/01-proto-sha256.mdand do exactly that task.
Then the next file in a new session. Or let tools/run-plan.sh docs/plans/M2a do that: it runs each
task in a fresh opencode run session with Laguna, and stops at the first task that does not end
with a commit, a clean tree and a done row. If a session ends with a stopped row in
docs/implementer-log.md, do not start the next task. Task 13 talks to straylight: Ornith must be
loadable and slot 0 should not be in heavy use while it runs.
For the reviewer: after task 13
git log --oneline master..m2a: thirteen commits with theImplemented-Bytrailer.- Copied files are unchanged:
for f in $(cd docs/plans/M2a/files && find . -type f); do cmp "docs/plans/M2a/files/$f" "$f"; done git diff master..m2a --stat -- docs/design.md docs/specs docs/plans AGENTS.md CLAUDE.md deny.tomlis empty.make gate,make audit,make verify-device.- Read every source file against its task and the spec. Probe from outside with inputs the tests do not contain, especially: malformed HTTP, a server that misbehaves mid-stream, several threads on the gate, and limits at their boundaries.
- Run the timing tests repeatedly under CPU load.
- Write findings under "Reviews" in
docs/implementer-log.md, and turn them into rows indocs/implementer-lessons.md, filling in "Seen again" for the M1 tips.