Add M2a design spec: the inference path
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,212 @@
|
||||
# M2a design: the inference path
|
||||
|
||||
Status: draft for owner review, 2026-09-17. M2 is split in two (`docs/decisions.md`). M2a is
|
||||
everything between `loopd` and `llama-server`: the forwarder, the HTTP and SSE client, the llama
|
||||
client with its timers and caps, a fake server for tests, and the startup self-test. M2b builds
|
||||
sessions, the turn loop and `bxctl chat` on top of it. Where this document and `docs/design.md`
|
||||
disagree, the brief wins. Measurements it relies on are in `docs/inference-contract.md`.
|
||||
|
||||
## 1. What M2a proves
|
||||
|
||||
| Claim | Checked by |
|
||||
|---|---|
|
||||
| `loopd` reaches the model only through a Unix socket | `inferproxy` tests; `loopd` has no TCP code |
|
||||
| A request is a faithful render of its inputs, and streamed deltas reassemble exactly | Fixture tests against transcripts recorded from straylight |
|
||||
| Silence, a dead server, a busy slot and a runaway thinking block each end in a defined way | Timing tests against the fake server |
|
||||
| A restart of the server under a request is survived | Retry tests; on device, the test kills and restarts its own `inferproxy` |
|
||||
| The deployed server is the one the config describes, tool calls parse, and turn 2 hits the cache | Startup self-test; `make verify-device` |
|
||||
|
||||
Out of scope: session logs, the turn loop, tools, the channel protocol, `bxctl chat`, the baseline
|
||||
token budget test (all M2b).
|
||||
|
||||
## 2. Components
|
||||
|
||||
```
|
||||
loopd ── infer.sock ── inferproxy ── TCP ── llama-server router (straylight:11434)
|
||||
```
|
||||
|
||||
### `inferproxy` (crate `inferproxy`)
|
||||
|
||||
`inferproxy --listen <socket path> --upstream <host:port>`. No config file. It logs nothing about
|
||||
traffic; it prints one line at start and one line per refused connection.
|
||||
|
||||
- Removes a stale socket file, binds, sets the socket file mode to 0600.
|
||||
- For each accepted connection: open one TCP connection upstream, copy bytes both ways on two
|
||||
threads, and pass a half-close on. When either side fails, close both.
|
||||
- At most 8 connections at a time. A token bucket on `accept`: burst 10, refill 2 per second. A
|
||||
connection over either limit is closed at once.
|
||||
- It never parses what it forwards.
|
||||
|
||||
### `proto::hash`
|
||||
|
||||
`pub fn sha256(data: &[u8]) -> Result<Hash32, HashError>`, a wrapper around the owner's `emsha`
|
||||
crate (1.0.4 or later). `proto`'s tests carry their own vectors: the empty message, `abc`, the
|
||||
million-`a` message, and messages of 55, 56, 63, 64 and 65 bytes, each also fed in two pieces.
|
||||
|
||||
### `loopd` modules
|
||||
|
||||
Each is one file under 500 lines with one purpose. None of them starts a thread.
|
||||
|
||||
| Module | Purpose |
|
||||
|---|---|
|
||||
| `config` | `config.toml` into a typed `Config`. Unknown keys are errors. Every limit in section 4 is a field with the default given there. |
|
||||
| `http` | HTTP/1.1 client over a `UnixStream`. One request per connection, `Connection: close`. Reads the status line and headers, then offers the body as a reader that handles chunked, content-length and read-to-close bodies. |
|
||||
| `sse` | Turns a body reader into a sequence of `data:` payloads. Handles lines split across reads, CRLF, comment lines and `[DONE]`. |
|
||||
| `llama::request` | Builds the chat-completions JSON from typed inputs: messages, tools, slot, sampling. |
|
||||
| `llama::assemble` | Folds streamed deltas into one assistant message, and keeps the latest per-chunk timings. |
|
||||
| `llama::chat` | One request from gate to completion: the waits, liveness, the thinking cap. |
|
||||
| `llama::retry` | `chat_with_retry`, and the slot gate. |
|
||||
| `llama::info` | `props`, `slots`, `tokenize`, and the cache-loss function. |
|
||||
| `selftest` | The three startup checks. `loopd selftest --config <path>` runs them and exits. |
|
||||
|
||||
`loopd` gains the dependencies `serde`, `serde_json` and `toml`, all already vetted.
|
||||
|
||||
## 3. Interfaces
|
||||
|
||||
```rust
|
||||
// What goes in. Messages are the four kinds the chat template knows.
|
||||
pub enum ChatMessage {
|
||||
System { content: String },
|
||||
User { content: String },
|
||||
Assistant { content: Option<String>, reasoning_content: Option<String>, tool_calls: Vec<proto::ToolCall> },
|
||||
Tool { tool_call_id: String, content: String },
|
||||
}
|
||||
pub struct ToolSchema { pub name: String, pub description: String, pub parameters: serde_json::Value }
|
||||
pub struct ChatRequest {
|
||||
pub slot: u32,
|
||||
pub messages: Vec<ChatMessage>,
|
||||
pub tools: Vec<ToolSchema>, // fixed for the epoch; may be empty
|
||||
pub thinking: bool,
|
||||
}
|
||||
|
||||
// What comes back while it runs.
|
||||
pub enum ChatEvent {
|
||||
Queued { ahead: usize }, // waiting for loopd's own slot gate
|
||||
Waiting { slot_busy: bool }, // sent, no byte yet; one event per /slots poll
|
||||
Progress { total: u64, cache: u64, processed: u64 },
|
||||
Reasoning(String),
|
||||
Content(String),
|
||||
ToolCallDelta { index: u32, id: Option<String>, name: Option<String>, arguments: String },
|
||||
ThinkingCapped { tokens: u64 },
|
||||
Retrying { attempt: u32, after_ms: u64, error: String },
|
||||
}
|
||||
|
||||
// What comes back at the end. Nothing is final before this.
|
||||
pub struct Completion {
|
||||
pub id: String,
|
||||
pub content: Option<String>,
|
||||
pub reasoning_content: Option<String>,
|
||||
pub tool_calls: Vec<proto::ToolCall>,
|
||||
pub finish_reason: FinishReason, // Stop | ToolCalls | Length
|
||||
pub timings: Timings, // cache_n, prompt_n, predicted_n
|
||||
pub reasoning_tokens: u64,
|
||||
pub thinking_capped: bool,
|
||||
}
|
||||
|
||||
pub enum InferError {
|
||||
Busy, // loopd's queue for the slot is full
|
||||
Connect(std::io::Error),
|
||||
WaitTimeout, // the slot stayed busy past the limit
|
||||
LoadTimeout, // the server stayed unavailable past the limit
|
||||
Stalled, // silence past the liveness limit
|
||||
StreamClosedEarly,
|
||||
ThinkingOverrun,
|
||||
Http { status: u16, body: String }, // body capped at 4 KiB
|
||||
Protocol(String),
|
||||
}
|
||||
|
||||
impl Client {
|
||||
pub fn chat(&self, req: &ChatRequest, on_event: &mut dyn FnMut(&ChatEvent)) -> Result<Completion, InferError>;
|
||||
pub fn chat_with_retry(&self, req: &ChatRequest, on_event: &mut dyn FnMut(&ChatEvent)) -> Result<Completion, InferError>;
|
||||
}
|
||||
```
|
||||
|
||||
`Assistant` uses the same three fields as `proto::LogRecord::Assistant`, so M2b can replay a log
|
||||
without conversion. Every request also carries, from config: the model id, `id_slot`,
|
||||
`cache_prompt: true`, `stream: true`, `return_progress: true`, `timings_per_token: true`,
|
||||
`reasoning_control: true`, `max_tokens`, and the sampling settings.
|
||||
|
||||
## 4. The life of one request
|
||||
|
||||
1. **Slot gate.** One request in flight per slot, inside `loopd`. Others queue, at most 8; beyond
|
||||
that the caller gets `Busy`. The gate is held for one request, never across a tool call or an
|
||||
approval, so a long tool does not block other sessions.
|
||||
2. **Send, then wait for the first byte.** Reads time out every 5 s (`poll_ms`). On each timeout
|
||||
the client asks `/slots` on a second connection and emits `Waiting`:
|
||||
- the slot is busy: keep waiting, up to 10 min (`busy_wait_ms`), then `WaitTimeout`
|
||||
- `/slots` fails or answers 503: the model is loading or the server is restarting; keep waiting,
|
||||
up to 3 min (`load_wait_ms`), then `LoadTimeout`
|
||||
- the slot is idle and still no byte for 30 s (`idle_grace_ms`): `Stalled`
|
||||
3. **After the first byte.** The read timeout becomes the liveness limit, 30 s (`liveness_ms`).
|
||||
Progress events and deltas both count. Silence past it is `Stalled`. There is no total deadline.
|
||||
4. **Thinking cap.** Reasoning tokens are counted from the per-chunk `predicted_n`. At 4,096
|
||||
(`thinking_cap`) the client posts `reasoning_end` on a second connection, emits
|
||||
`ThinkingCapped`, and keeps reading. If reasoning continues for 256 more tokens, or the control
|
||||
call does not return `success: true`, the client closes the connection, which cancels
|
||||
generation, and returns `ThinkingOverrun`. `max_tokens` is 8,192 as a backstop.
|
||||
5. **Done.** The final chunk carries `finish_reason`. The client returns the `Completion`. A stream
|
||||
that ends without one is `StreamClosedEarly`.
|
||||
|
||||
**Retry.** `chat_with_retry` retries `Connect`, `StreamClosedEarly`, `Stalled`, `LoadTimeout` and
|
||||
HTTP 503: 4 attempts in all, waiting 2 s, 8 s and 30 s with ±25% jitter, and never past 5 min in
|
||||
total. Each retry goes back through the slot gate and emits `Retrying`. Everything else is
|
||||
returned at once: `WaitTimeout` has already waited, HTTP 400 includes "the prompt does not fit",
|
||||
which M2b answers with compaction, and the rest are bugs. Retrying is safe because nothing is
|
||||
recorded until a completion is final, so a retry sends the same bytes again.
|
||||
|
||||
**Size caps in the client.** 16 KiB of headers, 1 MiB per SSE line, 4 MiB for a body that is not a
|
||||
stream. Exceeding one is `Protocol`.
|
||||
|
||||
**Cache loss.** `cache_outcome(previous: &Timings, current: &Timings) -> CacheOutcome` compares
|
||||
`current.cache_n` with the previous request's total (`cache_n + prompt_n + predicted_n`). More
|
||||
than 64 tokens short is a loss, reported with both numbers. It is a pure function; M2b logs it.
|
||||
|
||||
## 5. Startup self-test
|
||||
|
||||
`loopd` refuses to start unless all three pass. They run on the main slot and cost about 1,000
|
||||
tokens. If the router has unloaded the model, the first request loads it and the load wait covers
|
||||
that.
|
||||
|
||||
1. **The server is the one the config describes.** From `/props`: the SHA-256 of `chat_template`,
|
||||
per-slot `n_ctx` and `total_slots` equal the `[expect]` section of the config.
|
||||
2. **Tool calls parse.** One tool, a fixed prompt, and the completion must contain that tool call
|
||||
with valid JSON arguments and `finish_reason` `ToolCalls`.
|
||||
3. **Turn 2 hits the cache.** A two-turn exchange; `cache_outcome` for turn 2 must not be a loss.
|
||||
|
||||
## 6. Testing
|
||||
|
||||
**The fake server** is test support inside `loopd` (`tests/support/`), written by the design
|
||||
model. It listens on a Unix socket in a temporary directory and serves each connection from a
|
||||
script: a status, headers, and body pieces with a delay before each, ending in a clean close, a
|
||||
cut connection or silence. It records every request it receives, which M2b will use to check that
|
||||
each request extends the one before.
|
||||
|
||||
**Transcripts.** Bodies are real SSE transcripts recorded from straylight with synthetic prompts:
|
||||
a plain answer, thinking then an answer, a streamed tool call, progress events during a long
|
||||
prompt, a forced end to reasoning, and one raw chunked HTTP response for the `http` tests.
|
||||
|
||||
**Timing tests** set every limit to between 50 and 200 ms through `Config`, so none of them sleeps
|
||||
for a second.
|
||||
|
||||
**On device.** `make verify-device` builds the workspace, starts `inferproxy` against
|
||||
`straylight:11434`, runs `loopd selftest`, and then runs the ignored tests in
|
||||
`crates/loopd/tests/device.rs`: a capped thinking block, and a stream that survives its
|
||||
`inferproxy` being killed and restarted. The tests get the path of the `inferproxy` binary from
|
||||
the environment, so `loopd` does not depend on that crate. They use slot 0 and never touch the
|
||||
real `llama-server` process.
|
||||
|
||||
## 7. Threat model addition
|
||||
|
||||
A compromised `loopd` can degrade the shared `llama-server` for its other clients, by sending large
|
||||
prompts or requests that are not pinned to its own slots. This is accepted for v0. The harm is to
|
||||
availability only, the owner would see it, and preventing it would mean parsing untrusted input in
|
||||
`inferproxy`. The limits in `inferproxy` and the slot gate exist to stop accidents: a retry storm
|
||||
after a server restart, several sessions talking at once, a loop that spins.
|
||||
|
||||
## 8. How the work is handed over
|
||||
|
||||
As in M1 (`docs/specs/2026-09-17-pre-m1-design.md`, section 9), with the lessons in
|
||||
`docs/implementer-lessons.md` applied: an example for every kind of place a rule applies, tests
|
||||
that walk the data, and every syntactic form named. Tests, fixtures and the fake server are given;
|
||||
the implementer writes the code. Review happens once, after the last M2a task, before M2b is
|
||||
planned in detail.
|
||||
Reference in New Issue
Block a user