# M2a design: the inference path Status: draft for owner review, 2026-09-17. M2 is split in two (`docs/decisions.md`). M2a is everything between `loopd` and `llama-server`: the forwarder, the HTTP and SSE client, the llama client with its timers and caps, a fake server for tests, and the startup self-test. M2b builds sessions, the turn loop and `bxctl chat` on top of it. Where this document and `docs/design.md` disagree, the brief wins. Measurements it relies on are in `docs/inference-contract.md`. ## 1. What M2a proves | Claim | Checked by | |---|---| | `loopd` reaches the model only through a Unix socket | `inferproxy` tests; `loopd` has no TCP code | | A request is a faithful render of its inputs, and streamed deltas reassemble exactly | Fixture tests against transcripts recorded from straylight | | Silence, a dead server, a busy slot and a runaway thinking block each end in a defined way | Timing tests against the fake server | | A restart of the server under a request is survived | Retry tests; on device, the test kills and restarts its own `inferproxy` | | The deployed server is the one the config describes, tool calls parse, and turn 2 hits the cache | Startup self-test; `make verify-device` | Out of scope: session logs, the turn loop, tools, the channel protocol, `bxctl chat`, the baseline token budget test (all M2b). ## 2. Components ``` loopd ── infer.sock ── inferproxy ── TCP ── llama-server router (straylight:11434) ``` ### `inferproxy` (crate `inferproxy`) `inferproxy --listen --upstream `. No config file. It logs nothing about traffic; it prints one line at start and one line per refused connection. - Removes a stale socket file, binds, sets the socket file mode to 0600. - For each accepted connection: open one TCP connection upstream, copy bytes both ways on two threads, and pass a half-close on. When either side fails, close both. - At most 8 connections at a time. A token bucket on `accept`: burst 10, refill 2 per second. A connection over either limit is closed at once. - It never parses what it forwards. ### `proto::hash` `pub fn sha256(data: &[u8]) -> Result`, a wrapper around the owner's `emsha` crate (1.0.4 or later). `proto`'s tests carry their own vectors: the empty message, `abc`, the million-`a` message, and messages of 55, 56, 63, 64 and 65 bytes, each also fed in two pieces. ### `loopd` modules Each is one file under 500 lines with one purpose. None of them starts a thread. | Module | Purpose | |---|---| | `config` | `config.toml` into a typed `Config`. Unknown keys are errors. Every limit in section 4 is a field with the default given there. | | `http` | HTTP/1.1 client over a `UnixStream`. One request per connection, `Connection: close`. Reads the status line and headers, then offers the body as a reader that handles chunked, content-length and read-to-close bodies. | | `sse` | Turns a body reader into a sequence of `data:` payloads. Handles lines split across reads, CRLF, comment lines and `[DONE]`. | | `llama::request` | Builds the chat-completions JSON from typed inputs: messages, tools, slot, sampling. | | `llama::assemble` | Folds streamed deltas into one assistant message, and keeps the latest per-chunk timings. | | `llama::chat` | One request from gate to completion: the waits, liveness, the thinking cap. | | `llama::retry` | `chat_with_retry`, and the slot gate. | | `llama::info` | `props`, `slots`, `tokenize`, and the cache-loss function. | | `selftest` | The three startup checks. `loopd selftest --config ` runs them and exits. | `loopd` gains the dependencies `serde`, `serde_json` and `toml`, all already vetted. ## 3. Interfaces ```rust // What goes in. Messages are the four kinds the chat template knows. pub enum ChatMessage { System { content: String }, User { content: String }, Assistant { content: Option, reasoning_content: Option, tool_calls: Vec }, Tool { tool_call_id: String, content: String }, } pub struct ToolSchema { pub name: String, pub description: String, pub parameters: serde_json::Value } pub struct ChatRequest { pub slot: u32, pub messages: Vec, pub tools: Vec, // fixed for the epoch; may be empty pub thinking: bool, } // What comes back while it runs. pub enum ChatEvent { Queued { ahead: usize }, // waiting for loopd's own slot gate Waiting { slot_busy: bool }, // sent, no byte yet; one event per /slots poll Progress { total: u64, cache: u64, processed: u64 }, Reasoning(String), Content(String), ToolCallDelta { index: u32, id: Option, name: Option, arguments: String }, ThinkingCapped { tokens: u64 }, Retrying { attempt: u32, after_ms: u64, error: String }, } // What comes back at the end. Nothing is final before this. pub struct Completion { pub id: String, pub content: Option, pub reasoning_content: Option, pub tool_calls: Vec, pub finish_reason: FinishReason, // Stop | ToolCalls | Length pub timings: Timings, // cache_n, prompt_n, predicted_n pub reasoning_tokens: u64, pub thinking_capped: bool, } pub enum InferError { Busy, // loopd's queue for the slot is full Connect(std::io::Error), WaitTimeout, // the slot stayed busy past the limit LoadTimeout, // the server stayed unavailable past the limit Stalled, // silence past the liveness limit StreamClosedEarly, ThinkingOverrun, Http { status: u16, body: String }, // body capped at 4 KiB Protocol(String), } impl Client { pub fn chat(&self, req: &ChatRequest, on_event: &mut dyn FnMut(&ChatEvent)) -> Result; pub fn chat_with_retry(&self, req: &ChatRequest, on_event: &mut dyn FnMut(&ChatEvent)) -> Result; } ``` `Assistant` uses the same three fields as `proto::LogRecord::Assistant`, so M2b can replay a log without conversion. Every request also carries, from config: the model id, `id_slot`, `cache_prompt: true`, `stream: true`, `return_progress: true`, `timings_per_token: true`, `reasoning_control: true`, `max_tokens`, and the sampling settings. ## 4. The life of one request 1. **Slot gate.** One request in flight per slot, inside `loopd`. Others queue, at most 8; beyond that the caller gets `Busy`. The gate is held for one request, never across a tool call or an approval, so a long tool does not block other sessions. 2. **Send, then wait for the first byte.** Reads time out every 5 s (`poll_ms`). On each timeout the client asks `/slots` on a second connection and emits `Waiting`: - the slot is busy: keep waiting, up to 10 min (`busy_wait_ms`), then `WaitTimeout` - `/slots` fails or answers 503: the model is loading or the server is restarting; keep waiting, up to 3 min (`load_wait_ms`), then `LoadTimeout` - the slot is idle and still no byte for 30 s (`idle_grace_ms`): `Stalled` 3. **After the first byte.** The read timeout becomes the liveness limit, 30 s (`liveness_ms`). Progress events and deltas both count. Silence past it is `Stalled`. There is no total deadline. 4. **Thinking cap.** Reasoning tokens are counted from the per-chunk `predicted_n`. At 4,096 (`thinking_cap`) the client posts `reasoning_end` on a second connection, emits `ThinkingCapped`, and keeps reading. If reasoning continues for 256 more tokens, or the control call does not return `success: true`, the client closes the connection, which cancels generation, and returns `ThinkingOverrun`. `max_tokens` is 8,192 as a backstop. 5. **Done.** The final chunk carries `finish_reason`. The client returns the `Completion`. A stream that ends without one is `StreamClosedEarly`. **Retry.** `chat_with_retry` retries `Connect`, `StreamClosedEarly`, `Stalled`, `LoadTimeout` and HTTP 503: 4 attempts in all, waiting 2 s, 8 s and 30 s with ±25% jitter, and never past 5 min in total. Each retry goes back through the slot gate and emits `Retrying`. Everything else is returned at once: `WaitTimeout` has already waited, HTTP 400 includes "the prompt does not fit", which M2b answers with compaction, and the rest are bugs. Retrying is safe because nothing is recorded until a completion is final, so a retry sends the same bytes again. **Size caps in the client.** 16 KiB of headers, 1 MiB per SSE line, 4 MiB for a body that is not a stream. Exceeding one is `Protocol`. **Cache loss.** `cache_outcome(previous: &Timings, current: &Timings) -> CacheOutcome` compares `current.cache_n` with the previous request's total (`cache_n + prompt_n + predicted_n`). More than 64 tokens short is a loss, reported with both numbers. It is a pure function; M2b logs it. ## 5. Startup self-test `loopd` refuses to start unless all three pass. They run on the main slot and cost about 1,000 tokens. If the router has unloaded the model, the first request loads it and the load wait covers that. 1. **The server is the one the config describes.** From `/props`: the SHA-256 of `chat_template`, per-slot `n_ctx` and `total_slots` equal the `[expect]` section of the config. 2. **Tool calls parse.** One tool, a fixed prompt, and the completion must contain that tool call with valid JSON arguments and `finish_reason` `ToolCalls`. 3. **Turn 2 hits the cache.** A two-turn exchange; `cache_outcome` for turn 2 must not be a loss. ## 6. Testing **The fake server** is test support inside `loopd` (`tests/support/`), written by the design model. It listens on a Unix socket in a temporary directory and serves each connection from a script: a status, headers, and body pieces with a delay before each, ending in a clean close, a cut connection or silence. It records every request it receives, which M2b will use to check that each request extends the one before. **Transcripts.** Bodies are real SSE transcripts recorded from straylight with synthetic prompts: a plain answer, thinking then an answer, a streamed tool call, progress events during a long prompt, a forced end to reasoning, and one raw chunked HTTP response for the `http` tests. **Timing tests** set every limit to between 50 and 200 ms through `Config`, so none of them sleeps for a second. **On device.** `make verify-device` builds the workspace, starts `inferproxy` against `straylight:11434`, runs `loopd selftest`, and then runs the ignored tests in `crates/loopd/tests/device.rs`: a capped thinking block, and a stream that survives its `inferproxy` being killed and restarted. The tests get the path of the `inferproxy` binary from the environment, so `loopd` does not depend on that crate. They use slot 0 and never touch the real `llama-server` process. ## 7. Threat model addition A compromised `loopd` can degrade the shared `llama-server` for its other clients, by sending large prompts or requests that are not pinned to its own slots. This is accepted for v0. The harm is to availability only, the owner would see it, and preventing it would mean parsing untrusted input in `inferproxy`. The limits in `inferproxy` and the slot gate exist to stop accidents: a retry storm after a server restart, several sessions talking at once, a loop that spins. ## 8. How the work is handed over As in M1 (`docs/specs/2026-09-17-pre-m1-design.md`, section 9), with the lessons in `docs/implementer-lessons.md` applied: an example for every kind of place a rule applies, tests that walk the data, and every syntactic form named. Tests, fixtures and the fake server are given; the implementer writes the code. Review happens once, after the last M2a task, before M2b is planned in detail.