loopd's self-test caught the change (context per slot 131072 -> 262144, slots 2 -> 4). The device tests keep the expectation in one constant, and the M3a script matches it; verify-device passes 6 of 6 and the M3a device check passes. The inference contract notes which M0 findings rest on the old layout and need re-measuring. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
244 lines
14 KiB
Markdown
244 lines
14 KiB
Markdown
# Inference contract: M0 measurements
|
|
|
|
Measured 2026-09-17 against straylight by `spike/m0.py` (throwaway, Python stdlib). Requests went to
|
|
`https://straylight.scylla-hammerhead.ts.net:10000`, model id `ornith-1.5-35b-a3b`, through
|
|
`/v1/chat/completions` with `temperature 0.6, top_p 0.95, top_k 20`. Request fields were taken from
|
|
the server README at tag `b10809`, the build that is running.
|
|
|
|
**Conditions.** Every number below was taken while another session was generating on Ornith slot 1
|
|
(an 85k to 105k-token conversation at about 50 tokens/s). Token counts (`cache_n`, `prompt_n`) are
|
|
not affected by that. Throughput is, so section (a) reports a separate run taken after the GPU went
|
|
idle. Section (d) was also run with the GPU idle.
|
|
|
|
## Deployment change, 2026-09-20
|
|
|
|
Found in the M3a review on 2026-09-22, when `loopd`'s self-test refused the server
|
|
(`context per slot: expected 131072, got 262144`). The owner changed Ornith's preset in
|
|
`~/src/nixos/hw/straylight/default.nix` after measuring prompt-cache thrash between OpenCode, Hermes
|
|
and its subagents on two slots:
|
|
|
|
- `parallel = 4` with `kv-unified = true`: four slots share one 262,144-token KV pool. Any one
|
|
session may use the whole pool while the others are idle.
|
|
- A 16 GiB host prompt cache (was 8 GiB), `--models-max 3`, and a server-side
|
|
`reasoning-budget = 8192`.
|
|
|
|
Not re-measured yet. Two findings above rest on the old layout and need checking before M3b or M5
|
|
leans on them: (d), that a second session on another slot leaves the first slot's cache intact,
|
|
now that the slots share one pool; and P9's picture of eviction, where another client's long
|
|
prompt can now crowd a harness session out of the pool without touching its slot. The harness's
|
|
own expectations (`crates/loopd/tests/device.rs`, `tools/check-m3a-device.sh`) record the new
|
|
layout; `make verify-device` passes against it (6 of 6, 2026-09-22).
|
|
|
|
## What is running
|
|
|
|
| Item | Value |
|
|
|---|---|
|
|
| Build | llama.cpp `b10809-5266f24` (nixpkgs-unstable `llama-cpp-0.4.0`, Vulkan backend) |
|
|
| Mode | Router: one public endpoint, one child `llama-server` per model, `--models-max 2` |
|
|
| Public listener | `0.0.0.0:11434`, firewalled to the tailnet; Tailscale Serve adds HTTPS on `:10000` |
|
|
| Other clients | Open WebUI and OpenCode use the same endpoint and the same Ornith instance |
|
|
| Ornith flags | `--jinja --no-mmap --ctx-size 262144 --parallel 2 --cache-type-k q8_0 --cache-type-v q8_0 --flash-attn on --n-gpu-layers 999 --sleep-idle-seconds 21600 --hf-repo ornith-ai/Ornith-1.5-35B-A3B-GGUF:Q4_K_M` |
|
|
| Slots | Until 2026-09-20: 2, each `n_ctx` 131072 (the 262144 split, not shared). Since then: 4 over one unified 262144-token pool (`kv-unified`); `/props` reports `n_ctx` 262144 per slot. See "Deployment change, 2026-09-20" |
|
|
| Server default sampling | temperature 1.0, top_k 20, top_p 0.95, min_p 0.05. The harness must send its own. |
|
|
| Chat template | 7,828 bytes, sha256 `f55f52930aa8bf44ab5cb85f99370fcc3c56e9a85640b812086d5330bce5d86b` |
|
|
| Source of truth for flags | `~/src/nixos/hw/straylight/default.nix` on straylight, not this repo |
|
|
|
|
Differences from the design brief: KV cache is q8_0, not f16. There are two slots, not three. The
|
|
server is shared, so slots are not reserved for the harness. Weights and KV cache are dropped after
|
|
six idle hours. Host memory was 110 of 125 GB in use with Laguna S 2.1 and Ornith both loaded.
|
|
|
|
## Findings
|
|
|
|
### (b) Cache reuse over a 3-turn conversation: passes
|
|
|
|
Four tool schemas, thinking on, `reasoning_content` and `tool_calls` echoed back exactly as received.
|
|
|
|
| Request | `cache_n` | `prompt_n` |
|
|
|---|---|---|
|
|
| turn 1, request 1 | 0 | 539 |
|
|
| turn 1, request 2 (after tool result) | 591 | 29 |
|
|
| turn 2, request 1 | 680 | 27 |
|
|
| turn 2, request 2 | 759 | 40 |
|
|
| turn 3, request 1 | 840 | 26 |
|
|
|
|
Each request processes only its new tokens. Raw timing fields for one request:
|
|
`{"cache_n": 591, "prompt_n": 29, "prompt_ms": 227.163, "prompt_per_second": 127.66, "predicted_n": 61, "predicted_ms": 1762.327, "predicted_per_second": 34.05}`.
|
|
|
|
The whole baseline here (system line, four tool schemas, first user message) was 539 tokens, so the
|
|
3,000-token baseline budget is realistic.
|
|
|
|
### (c) Tool-call parsing through chat-completions: 0 failures in 20
|
|
|
|
Five prompts for each of four tools. Every response had `finish_reason: "tool_calls"`, the expected
|
|
tool, valid JSON arguments and all required arguments. Types survived: an integer `timeout_s`, a
|
|
nested `headers` object, and strings containing quotes, `&`, `<>` and embedded JSON.
|
|
|
|
### (e) Thinking blocks and the cache
|
|
|
|
- Ornith's template renders the `<think>` block of **every** assistant turn, not only the last one.
|
|
It does not strip earlier reasoning. The brief's "known risk" does not apply as long as the
|
|
harness sends `reasoning_content` back unchanged.
|
|
- If the harness drops `reasoning_content`, the prompt diverges at the latest assistant turn. The
|
|
cost was small (`prompt_n` 65 to 80 instead of 27 to 40), because the server keeps a checkpoint
|
|
near the end of the previous request.
|
|
- A change anywhere earlier costs a full re-read. Editing turn 3 or turn 2 of a 5-turn, 7.6k-token
|
|
conversation gave `cache_n` 34. Changing text 10k tokens into a 24k-token prompt gave `cache_n` 0
|
|
and 26 s of prompt processing. This confirms the brief: no partial rewind in practice.
|
|
- The server sometimes restored an older prompt from its host-RAM prompt cache (`--cache-ram`,
|
|
default 8 GiB): resending the original 24k prompt after the edited one gave `cache_n` 23758. It
|
|
did not do so every time. Do not design around it.
|
|
|
|
### (h) Changing the tool list mid-session: full invalidation
|
|
|
|
The template renders tool schemas at the very top of the prompt, before the system text. Adding a
|
|
fifth tool at turn 3 gave `cache_n` 23, `prompt_n` 920. The `tools` array must be fixed for the
|
|
whole epoch.
|
|
|
|
Progressive disclosure still works if the schema arrives as a tool result:
|
|
|
|
| Variant | Result |
|
|
|---|---|
|
|
| `find_tool` returns a schema, model calls it through a fixed `call_tool(name, arguments)` meta-tool | 4 of 5 correct. The one miss called `call_tool` without `find_tool` first, which `brokerd` can reject. |
|
|
| `find_tool` returns a schema, model calls the new tool directly by name | 0 of 5. The server's grammar only allows declared names, so the model was forced into a **wrong declared tool**: three times it emitted `write_file` with placeholder content. |
|
|
|
|
The second row is a safety finding, not only a cache one. Never tell the model to call a tool that
|
|
is not in the `tools` array.
|
|
|
|
### (i) Liveness during prompt processing
|
|
|
|
With `stream: true` and `return_progress: true`, a 16k-token prefill produced 11 `prompt_progress`
|
|
events and the longest silence was 2.2 s. Without `return_progress` the stream is silent for the
|
|
whole prefill. The liveness timeout in the brief needs this field.
|
|
|
|
### (g) Unix socket, co-location
|
|
|
|
`--host` accepts a path ending in `.sock` (README, build b10809). The harness and `llama-server`
|
|
are on the same host. In router mode the router sets each child's host and port itself, so only the
|
|
router's public listener could move to a socket, and Open WebUI, OpenCode and Tailscale Serve need
|
|
it on TCP. `loopd` runs with `--network=none` and cannot reach host loopback. So `inferproxy` stays.
|
|
|
|
### (j) What a queued request receives: nothing
|
|
|
|
Measured 2026-09-17 before the M2 design. Request A generated on slot 0 for 10 s. Request B, pinned
|
|
to the same slot with `stream` and `return_progress`, was sent while A was running. B received no
|
|
bytes at all, not even response headers, until A finished: first byte at 10.02 s, then progress
|
|
events and tokens as usual. The stream cannot tell "queued" from "dead". `GET /slots?model=…` can:
|
|
it shows `is_processing` for the slot. So the wait before the first byte needs its own, longer
|
|
limit, and `loopd` can poll `/slots` during it to tell a busy slot from a dead server.
|
|
|
|
### (k) Runaway control: `reasoning_control` works, and costs one turn of cache
|
|
|
|
`--reasoning-budget` is a server flag, not a request field. Per request there is `max_tokens`, and
|
|
`reasoning_control: true` plus `POST /v1/chat/completions/control` with the completion's `id`,
|
|
`action: "reasoning_end"` and `model`.
|
|
|
|
Measured 2026-09-17: the control call was sent after 150 streamed reasoning chunks and returned
|
|
`{"success": true}`. Three more reasoning chunks arrived, then the model wrote its answer and
|
|
finished with `finish_reason: "stop"`.
|
|
|
|
The next turn, with the capped turn replayed exactly as streamed, had `cache_n` 85 and `prompt_n`
|
|
743: the server re-read the whole capped assistant turn. The same two turns with a natural end to
|
|
thinking gave `cache_n` 167, `prompt_n` 22. So what the server generated at a forced end is not
|
|
what the template renders from the replayed message, and the hybrid cache falls back to the
|
|
checkpoint at the end of the previous prompt. The cost is bounded by the thinking cap plus one
|
|
answer, and it is paid once. The control run also shows that streamed deltas, concatenated, replay
|
|
byte for byte.
|
|
|
|
### (l) Closing the connection cancels generation
|
|
|
|
Measured 2026-09-17: a streaming request was closed by the client while the model was thinking.
|
|
`/slots` showed the slot idle 0.1 s later. A client that aborts frees the GPU.
|
|
|
|
### (m) Exact token counts while streaming
|
|
|
|
With `timings_per_token: true` every streamed chunk carries a `timings` object, including
|
|
`predicted_n`, the number of tokens generated so far. A thinking cap can count real tokens.
|
|
|
|
### (n) Tool-call arguments stream token by token
|
|
|
|
Measured 2026-09-17: the model wrote a 1,500-word essay into a `write_file` call. 1,999 tokens over
|
|
30 s arrived as 1,948 tool-call deltas, the longest gap between events was 0.33 s, and the
|
|
concatenated `arguments` fragments were valid JSON. The server does not hold a tool call back
|
|
until it parses, so a silence-based liveness limit is safe during long tool calls.
|
|
|
|
### (a) Throughput
|
|
|
|
Method: timed requests, server-reported `timings`, thinking off. "At depth 32k" means a 2k-token
|
|
suffix appended to a cached 32k-token prefix, then generation from there.
|
|
|
|
| Measurement | GPU otherwise idle | Other Ornith slot generating |
|
|
|---|---|---|
|
|
| Prompt processing at depth 0 (2k tokens) | 1,170 tokens/s | 875 to 1,040 tokens/s |
|
|
| Prompt processing, average over 0 to 32k | 898 tokens/s (32k in 36 s) | 929 to 936 tokens/s over 0 to 24k |
|
|
| Prompt processing at depth 32k (2k suffix) | 669 tokens/s | not measured |
|
|
| Generation at depth 0 | 69 tokens/s | 28 to 34 tokens/s |
|
|
| Generation at depth 32k | 60 tokens/s | not measured |
|
|
|
|
The idle numbers match the brief's reference figures (about 1,100 and 700 tokens/s, 60 falling to
|
|
49). A second active session on the same model roughly halves generation speed.
|
|
|
|
### (d) Slot pinning and eviction
|
|
|
|
Run with the GPU otherwise idle and both slots free.
|
|
|
|
**Pinning works.** Session A pinned to slot 0 and session B pinned to slot 1 (6k tokens each) did not
|
|
disturb each other: A's second turn had `cache_n` 6028, `prompt_n` 17.
|
|
|
|
**Eviction is recovered from the server's host-RAM prompt cache** (`--cache-ram`, default 8 GiB,
|
|
with `--cache-idle-slots` on by default), as long as the session comes back on the same slot.
|
|
|
|
| Scenario | Result for the returning session |
|
|
|---|---|
|
|
| Two unpinned 3k-token requests from "other clients" between turns of A and B | A: `cache_n` 6088, `prompt_n` 18. B: `cache_n` 6107, `prompt_n` 17. |
|
|
| A (30k tokens) on slot 0, then three unrelated 8k prompts pinned to slot 0, then A again on slot 0 | `cache_n` 30695, `prompt_n` 16, 0.43 s |
|
|
| The same session A then sent to slot 1 instead | `cache_n` 0, `prompt_n` 30728, 34.5 s |
|
|
| Two sessions sharing a 3k baseline (same system text and tools), 10k tokens each, taking turns on slot 0 for three rounds | Every switch after the first: `cache_n` about 13k, `prompt_n` about 320, 1 s. The second session's first request also reused the 3,031-token shared baseline. |
|
|
|
|
Consequences:
|
|
|
|
- A session must always use the same slot. Moving it to another slot costs a full re-read.
|
|
- Several Mattermost threads can share the main slot. Switching between them costs about a second,
|
|
not a re-read, while their saved states fit in the prompt cache.
|
|
- Another client taking a harness slot is usually harmless for the same reason.
|
|
- This does not survive a model unload (`--models-max 2`, or six idle hours), and the 8 GiB limit
|
|
was not probed. The size of one saved session was not measured. `loopd` still has to treat a cold
|
|
cache as a normal event.
|
|
- One exception was seen in (e): after a prompt that shared a long prefix with the saved one but
|
|
differed in the middle, the saved prompt was restored on one attempt and not on another.
|
|
|
|
## Recommendation on the open question
|
|
|
|
Use the server's chat-completions endpoint with server-side tool parsing. Do not render the
|
|
template in-process. The cache measurements pass and tool parsing had no failures. The conditions
|
|
are:
|
|
|
|
1. The session log stores each assistant message exactly as returned (`content`,
|
|
`reasoning_content`, `tool_calls`) and replays it unchanged.
|
|
2. The `tools` array is fixed per epoch. Tools outside the core set are reached through `find_tool`
|
|
and a `call_tool` meta-tool.
|
|
3. Every request carries `id_slot`, `cache_prompt: true`, the sampling settings, `stream: true` and
|
|
`return_progress: true`.
|
|
4. The startup self-test compares the template hash and `n_ctx` from `/props?model=...` with the
|
|
values recorded here.
|
|
|
|
## Serving setup
|
|
|
|
Decided 2026-09-17: Boxmaker uses the shared router and the shared Ornith instance
|
|
(`docs/decisions.md`). A dedicated `llama-server` on a Unix socket was considered. It would have
|
|
given reserved slots and removed `inferproxy`, but a second resident copy of Ornith (about 22 GB
|
|
plus KV cache) does not fit beside Laguna's 69 GB under the 104 GiB GPU memory cap.
|
|
|
|
What the shared instance means for the harness:
|
|
|
|
- `id_slot` pins the harness's own requests. It does not reserve the slot. An unpinned request
|
|
from another client can take it, but the evicted session is normally restored from the host-RAM
|
|
prompt cache (see (d)). Loading a third model can unload Ornith (`--models-max 2`), and six idle
|
|
hours drop weights and cache; neither is recoverable. `loopd` must treat a cold cache as a
|
|
normal event.
|
|
- A pinned request can wait behind another client's generation on the same slot.
|
|
- Another active session on Ornith roughly halves generation speed.
|
|
- Serving flags are owned by `~/src/nixos`. This repo records expected values and checks them at
|
|
startup.
|
|
- Suggested change on the NixOS side: Ornith `parallel = 4`, `ctx-size = 524288`, so the harness
|
|
can pin three slots and one is left for other clients. Extra KV memory not yet measured.
|