The tasks build the inference path: emsha-backed SHA-256, inferproxy, config, a hand-written HTTP and SSE client, request building, delta assembly, the chat state machine, the thinking cap, the slot gate with retry, the startup self-test and on-device verification. Everything the tasks copy in was checked against a private reference implementation: the gate passes after each task in order, the timing tests pass repeatedly under CPU load, and the reference passes the self-test and all four device checks on straylight. Expected results for the recorded streams were derived by a separate script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
3.6 KiB
M2a task 10: the thinking cap
Branch: m2a (run git switch m2a; git status --short must be empty, otherwise stop)
Commit subject: Add the thinking cap to Client::chat
Goal
Stop a model that thinks for too long. When a completion's reasoning reaches thinking_cap tokens,
tell the server to end the reasoning block; the model then writes its answer. If it keeps thinking
anyway, give up on the request.
Context
The server has a control call for a running completion, measured to work: three more reasoning chunks arrived after it, then the answer.
POST /v1/chat/completions/control
{"id":"<the completion id from the stream>","action":"reasoning_end","model":"<cfg.infer.model>"}
→ {"success":true}
It is made on a second connection while the stream is still open on the first. No thread is needed: between two reads of the stream, make the call, then go on reading. The stream's bytes wait in the socket meanwhile.
Files
- Copy:
crates/loopd/tests/cap.rs - Modify:
crates/loopd/src/llama/chat.rs,crates/loopd/src/llama/info.rs,docs/implementer-log.md
Interfaces
Add to info.rs:
impl Client {
/// Tells the server to end the reasoning block of a running completion. Ok(success).
pub(crate) fn end_reasoning(&self, completion_id: &str) -> Result<bool, InferError>;
}
It goes through call, so a status other than 200 is already an Err. A 200 body without a
boolean success is Protocol. This is the server's format: ignore other fields (message).
Change to step 4 of chat
After pushing a chunk into the assembler and passing its events on, if
assembler.in_reasoning(), let tokens = assembler.reasoning_tokens() and:
| Cap already fired? | Condition | Do |
|---|---|---|
| no | tokens >= thinking_cap |
call end_reasoning(id). If it is Ok(true): remember tokens as the point where the cap fired, and emit ChatEvent::ThinkingCapped { tokens }. For Ok(false) or any Err: return ThinkingOverrun. |
yes, at at |
tokens >= at + thinking_overrun |
return ThinkingOverrun |
| either | otherwise | nothing |
Make the call once. After the cap has fired, later reasoning chunks must not trigger it again.
id is assembler.id(); every chunk of the real server carries it.
In step 5, pass whether the cap fired: assembler.finish(cap_fired). Completion::thinking_capped
tells the caller, and reasoning_tokens keeps counting after the cap, so the caller can see how
far the model went.
When nothing in a completion is reasoning, none of this happens, whatever the cap.
Steps
- 1. Copy.
git switch m2a, thencp docs/plans/M2a/files/crates/loopd/tests/cap.rs crates/loopd/tests/. The recording it uses,capped.http, has 63 reasoning chunks and then an answer. The fake server plays all of it whatever the client does, so the tests choosethinking_capandthinking_overrunto reach each row of the table. - 2. See the test fail.
cargo test -p loopd --test cap. Expected: it compiles, and 4 of 6 tests fail, because nothing caps anything yet. - 3. Make the change. Run
cargo fmt --all. - 4. See the tests pass.
cargo test -p loopd --test cap --test chat. Expected:6 passedand13 passed. - 5. Run the gate.
make gate. Expected last line:gate: ok. - 6. Log and commit.
git add crates/loopd docs/implementer-log.md && git commit
Done when
cargo test -p loopd --test capreports 6 passed and--test chatstill 13;make gateprintsgate: ok.
Stop and report if
the_allowance_is_exactcannot pass with the comparison written as>=in both rows.