# M2a task 10: the thinking cap **Branch:** `m2a` (run `git switch m2a`; `git status --short` must be empty, otherwise stop) **Commit subject:** `Add the thinking cap to Client::chat` ## Goal Stop a model that thinks for too long. When a completion's reasoning reaches `thinking_cap` tokens, tell the server to end the reasoning block; the model then writes its answer. If it keeps thinking anyway, give up on the request. ## Context The server has a control call for a running completion, measured to work: three more reasoning chunks arrived after it, then the answer. ``` POST /v1/chat/completions/control {"id":"","action":"reasoning_end","model":""} → {"success":true} ``` It is made on a **second connection** while the stream is still open on the first. No thread is needed: between two reads of the stream, make the call, then go on reading. The stream's bytes wait in the socket meanwhile. ## Files - Copy: `crates/loopd/tests/cap.rs` - Modify: `crates/loopd/src/llama/chat.rs`, `crates/loopd/src/llama/info.rs`, `docs/implementer-log.md` ## Interfaces Add to `info.rs`: ```rust impl Client { /// Tells the server to end the reasoning block of a running completion. Ok(success). pub(crate) fn end_reasoning(&self, completion_id: &str) -> Result; } ``` It goes through `call`, so a status other than 200 is already an `Err`. A 200 body without a boolean `success` is `Protocol`. This is the server's format: ignore other fields (`message`). ## Change to step 4 of `chat` After pushing a chunk into the assembler and passing its events on, if `assembler.in_reasoning()`, let `tokens = assembler.reasoning_tokens()` and: | Cap already fired? | Condition | Do | |---|---|---| | no | `tokens >= thinking_cap` | call `end_reasoning(id)`. If it is `Ok(true)`: remember `tokens` as the point where the cap fired, and emit `ChatEvent::ThinkingCapped { tokens }`. For `Ok(false)` or any `Err`: return `ThinkingOverrun`. | | yes, at `at` | `tokens >= at + thinking_overrun` | return `ThinkingOverrun` | | either | otherwise | nothing | Make the call **once**. After the cap has fired, later reasoning chunks must not trigger it again. `id` is `assembler.id()`; every chunk of the real server carries it. In step 5, pass whether the cap fired: `assembler.finish(cap_fired)`. `Completion::thinking_capped` tells the caller, and `reasoning_tokens` keeps counting after the cap, so the caller can see how far the model went. When nothing in a completion is reasoning, none of this happens, whatever the cap. ## Steps - [ ] **1. Copy.** `git switch m2a`, then `cp docs/plans/M2a/files/crates/loopd/tests/cap.rs crates/loopd/tests/`. The recording it uses, `capped.http`, has 63 reasoning chunks and then an answer. The fake server plays all of it whatever the client does, so the tests choose `thinking_cap` and `thinking_overrun` to reach each row of the table. - [ ] **2. See the test fail.** `cargo test -p loopd --test cap`. Expected: it compiles, and 4 of 6 tests fail, because nothing caps anything yet. - [ ] **3. Make the change.** Run `cargo fmt --all`. - [ ] **4. See the tests pass.** `cargo test -p loopd --test cap --test chat`. Expected: `6 passed` and `13 passed`. - [ ] **5. Run the gate.** `make gate`. Expected last line: `gate: ok`. - [ ] **6. Log and commit.** `git add crates/loopd docs/implementer-log.md && git commit` ## Done when - `cargo test -p loopd --test cap` reports 6 passed and `--test chat` still 13; `make gate` prints `gate: ok`. ## Stop and report if - `the_allowance_is_exact` cannot pass with the comparison written as `>=` in both rows.