M4 is split into M4a and M4b. The M4a spec rests on facts checked against the owner's server and Mattermost's source at v11.11.0: the REST and WebSocket shapes, and that clients will not post a message starting with '/'. Records the design decisions, proposes P15 (secrets from a systemd credential, the environment or an owner-only file), and adds the run-time rows egress.md was missing since M2a and M3b. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
7.2 KiB
7.2 KiB
Boxmaker milestones and session prompts
One session per milestone. The design brief is docs/design.md.
Milestones
| # | Milestone | Proves | Verified by |
|---|---|---|---|
| M0 | Measurement spike | The cache and tool-call assumptions hold on straylight | Numbers in docs/inference-contract.md |
| M1 | Workspace, proto, gate |
The type-level authority model compiles | make gate |
| M2 | loopd core + bxctl chat |
Append-only loop, baseline ≤ 3k tokens, self-test, runaway control, with one fake in-process tool | Prefix-extension property test; turn-2 cache hit on straylight |
| M3a | brokerd decisions |
Grants, deny by default, taint, hash-chained audit, ask approvals through bxctl, two sockets, loopd's broker port; tools do not run yet |
Tests for every deny path; audit tampering suite; a denial reaches Ornith on straylight as a tool result |
| M3b | brokerd runner |
Rootless container runner, four tools (read file, write file, sandboxed shell, HTTP fetch with host allowlist), egress proxy | A container with no network grant cannot reach the network; fetch refused for hosts outside the grant |
| M4a | gatewayd + Mattermost conversations |
Bot connection over TLS, user-ID allowlist, threads as sessions, bursts as one turn, typing while working, catching up after downtime, the secret store, no listening port | A direct message answered on the real server; a non-allowlisted user gets no response |
| M4b | Approvals over Mattermost | gateway.sock, approval posts, answers by reaction or !approve/!deny, long approvals in parts |
An ask call approved and refused from Mattermost; the audit log names the post |
| M5 | Scheduler + memory | Heartbeat and cron on their own slot in isolated sessions; core block, recall and remember with provenance, FTS5 index, flush-then-compact at idle |
Scheduled run leaves the main slot's cache intact; a fact from epoch 1 is recalled in epoch 2; a note written after reading untrusted content is marked on recall |
| M5b | Embedding index (when justified) | Local embedding server, hybrid retrieval behind the Retriever trait, rebuild on model change |
A paraphrased query finds a note that lexical search misses; no memory text leaves the host |
| M6 | Subagents + cloud consult | Subagents on their slot returning summaries; consult tool gated by data-class grants with full payload log | A tainted session is refused; payload log matches what was sent |
| M7 | Split deployment | Each role in its own rootless container; loopd with no network; socket volume permissions; egress proxy for tools; host ACL policy in deploy/ |
From inside the loopd container, requests to the internet, the tailnet and the host all fail; a tool container without a host grant cannot reach the tailnet |
Per-milestone prompt template
Read AGENTS.md, docs/design.md and docs/inference-contract.md. This session delivers milestone
M<n>: <one-line goal from the table>.
Scope: <bullets>. Out of scope: everything else, including later milestones.
Start by restating the acceptance checks from the milestone table as tests or scripted checks.
Write a plan, show it to me, then execute it with subagents, tests first. Verify external APIs
from their documentation before use. Finish with `make gate` and the on-device verification
named in the table; report exit status, last lines, and what you observed on straylight.
If the design brief turns out to be wrong, stop and propose the change to docs/design.md as its
own commit before building on it.
Session 1 prompt (M0 spike + M1 skeleton)
As written in the kickoff pack, kept for the record. It is superseded: M0 is done
(docs/inference-contract.md), and M1 is planned in docs/plans/M1/ from
docs/specs/2026-09-17-pre-m1-design.md.
```text
You are starting a new Rust project: Boxmaker, a sovereign personal agent harness. Read docs/design.md
first; it is the design brief and it is binding. If anything in it looks wrong or contradicts
what you measure, stop and tell me rather than working around it.
This session has two deliverables, in order.
DELIVERABLE 1 — M0 measurement spike (throwaway code, kept findings)
The inference host is `straylight` (Strix Halo, llama.cpp llama-server, Ornith-1.5-35B-A3B,
~128k context). Before any architecture is built, measure what the design assumes. Ask me for
the server URL and how you can reach it; if you cannot reach it, stop and say exactly what is
missing.
Fetch the llama-server README for the build running on straylight and the model card before
using any endpoint or parameter. Do not rely on memory for request fields. Then measure and
record in docs/inference-contract.md:
a. Prompt-processing and generation throughput at depth 0 and ~32k (llama-bench or timed
requests; say which).
b. A 3-turn conversation with 4 tool schemas and thinking enabled: tokens processed versus
tokens reused from cache on each turn. Show the raw timing fields.
c. Whether tool calls come back correctly parsed through the chat-completions endpoint with
the model's shipped template. Run 20 trials across 4 tools and report the failure count.
d. Whether requests can be pinned to a slot, and whether a second session on another slot
leaves the first slot's cache intact.
e. What happens to the cache when an earlier assistant turn contained a thinking block.
f. The exact launch flags in use, and the flags you recommend, with reasons.
g. Whether llama-server can listen on a Unix socket, and whether it runs on the same host as
the harness containers. This decides whether `inferproxy` is needed.
Spike code goes in spike/ and is labelled throwaway. Conclude with a recommendation on the open
question in the brief: server-side chat-completions, or in-process template rendering.
Stop after M0 and show me the findings before starting M1.
DELIVERABLE 2 — M1 workspace skeleton
- AGENTS.md at the repo root: project standards drawn from docs/design.md "Code constraints",
plus how to run the gate and how to verify against straylight.
- Cargo workspace with the crates named in the brief. `proto` gets the real shared types:
session log records, tool request and result, data-class labels, grant, decision, audit
record. Other crates are minimal binaries that compile.
- Makefile with `make gate`: cargo fmt --check, clippy with warnings denied, cargo test,
cargo-deny, and a check that fails on any source file over 500 lines.
- docs/dependencies.md and docs/egress.md, initially short and accurate.
- Unit tests for `proto` serialization round trips and for the rule that a `Decision` cannot
be constructed outside the policy module (a compile-fail test is acceptable).
Working rules
- Verify every external crate's current API on docs.rs before using it, and check its
maintenance status and license. If you cannot fetch documentation, say so; do not guess.
- Write a short plan before M1 and execute it with subagents. Tests first.
- Run `make gate` before declaring anything done. Report the exit status and last lines.
- One logical change per commit. Never commit runtime data, secrets, or spike output that
contains conversation content.
- Keep it small. If a feature is not in the brief, do not add it. Propose it instead.