| 2026-09-18 |
M3 is split. M3a is the decision path: grants and matching, session taint, the hash-chained audit log, approvals through bxctl, the broker.sock and admin.sock protocol, and loopd's BrokerPort; tools do not run (the runner is a trait with a refusing implementation). M3b is the Podman runner, the four tools and the egress proxy. Review happens after each. |
The security logic is reviewed before real tools run on it, as with M2a and M2b. Spec docs/specs/2026-09-18-m3a-decision-path.md. |
| 2026-09-18 |
Grant loading fails closed: one invalid grant file denies every call (grants_invalid) until it is fixed. M3 grants may not set secret or patterns; each tool has fixed rules for which constraints it takes. Among matching grants the most restrictive mode wins (deny, ask, auto). An approval re-decides against the current grants and taint, and runs only if the outcome is still ask or auto. |
A skipped, mistyped deny grant would silently become an allow. Pattern-matching shell commands is a false comfort; the container is the boundary. |
| 2026-09-18 |
Every fail-closed state has an entry in docs/runbook.md with what you see, why, how to confirm, how to fix and how to check; its message ends with see docs/runbook.md#<entry>, and a gate script checks that every referenced entry exists. |
Owner's requirement: a refusing system must come with clear, actionable remediation. |
| 2026-09-18 |
brokerd listens on two sockets: broker.sock (tool requests only, for loopd) and admin.sock (approvals and grant checks, for bxctl). Each refuses the other's messages with forbidden. Replaces the single broker.sock of the pre-M1 spec. |
Anything that could reach the approving socket could approve its own calls. Until M7 all roles run as the owner's user, so the split is enforced in code; M7 enforces it with mounts. |
| 2026-09-18 |
The audit log records events (decision, approval, result, recovery, accepted_break) in one chain. A decision is on disk before anything runs; results are recorded by hash and size, not content. A broken chain stops brokerd until the owner runs it once with --accept-break, which records the break; nothing is ever repaired or deleted. Chain verification is a pure function in proto shared by brokerd and bxctl audit verify. Approval ids are the seq of the decision record. |
A secret read must not be copied into the audit log. Verification must work when brokerd refuses to start. |
| 2026-09-18 |
Secrets move to M4. In M3 a grant that sets secret is invalid. |
None of M3's tools needs a secret; the first is the Mattermost bot token. |
| 2026-09-18 |
Until M7, brokerd and its containers run as the owner's user through rootless Podman; a container escape is the owner's user. Containers are hardened (--network=none unless granted, --read-only, --cap-drop=all, no-new-privileges, process and memory limits). |
Separate users per role are M7's work. Measured on straylight: rootless Podman 5.8.6 with crun, 40 to 80 ms per container. |
| 2026-09-18 |
Tool containers run from one OCI image built by Nix (dockerTools) holding static toolkit, busybox, curl and the CA bundle, loaded with podman load and named by digest. No registry pull at call time. The owner's Gitea registry is the route when there is more than one host. |
Pinned contents with no egress. Nix already builds the three binaries statically. |
| 2026-09-18 |
http_fetch runs curl in a container with no network, through a per-call SOCKS5 proxy in its own container on a mounted Unix socket. The proxy is ours (in toolkit), accepts host names only, checks each against the grant's hosts, and copies bytes; TLS stays end to end in the tool container. clock moves into loopd. |
No TLS stack of our own. Redirects are checked at the proxy. Measured on straylight 2026-09-18: allowed host 200 in 0.26 s; other host and IP literal refused. |
| 2026-09-18 |
M2b: the baseline (system prompt, core tool schemas, memory/core.md) is snapshotted per epoch into sessions/<id>/<epoch>.baseline.json; resume uses the snapshot, so edits apply only to later sessions. The agent is called Boxmaker; the first system.md is drafted by the design model and edited by the owner. A Usage log record follows each Assistant record. bxctl chat shows reasoning dimmed by default. The core tool set is clock, find_tool, call_tool; echo is the first discoverable tool. |
Owner's choices during the M2b design review; spec docs/specs/2026-09-18-m2b-agent-loop.md. Dimmed thinking shows progress and helps spot stalls. |
| 2026-09-17 |
A compromised loopd can degrade the shared llama-server for other clients (large prompts, unpinned requests). Accepted for v0 and written into the threat model. inferproxy stays a byte forwarder with a connection cap and an accept-rate limit; it does not enforce slot or model policy. If isolation is wanted later, do it on the server side. |
The harm is availability only, the owner would notice, and a policy proxy would put a parser for untrusted input into the component meant to have none. |
| 2026-09-17 |
M2a limits: 10 min wait for a busy slot, 3 min for a model load, 30 s liveness after the first byte, thinking cap 4,096 tokens with max_tokens 8,192 as backstop. All are config values. The slot gate is held per request, never across a tool call. |
Owner's choices during the M2a design review; measurements (j) to (n) in docs/inference-contract.md. |
| 2026-09-17 |
SHA-256 comes from the owner's emsha crate, version 1.0.4 or later, not sha2 and not hand-written code in this repo. It is wrapped behind one function in proto, whose tests carry their own vectors: abc, the million-a message, and lengths 55, 56, 63, 64 and 65. The crate's custom licence is not an issue: the owner is its author. |
No dependencies, no unsafe, no_std, owner-maintained. Version 1.0.3 hashed every message of length 63 mod 64 wrongly; a differential test against sha256sum found it while the crate was being vetted, and the owner fixed it in 1.0.4 (3,204 cases pass). |
| 2026-09-17 |
M2 is split. M2a is the inference path (inferproxy, HTTP and SSE client, llama client, fake server, startup self-test). M2b is the agent loop (session log, turn loop, channel protocol, bxctl chat). Review happens after each. |
The risky, timing-dependent work gets reviewed before the turn loop is built on it. |
| 2026-09-17 |
Tool results are untrusted by default: each grant has untrusted, default true, and brokerd tracks a per-session untrusted flag beside taint. Review of Laguna's work is once per milestone, as an experiment for M1 to be revisited after the M1 review. Laguna keeps docs/implementer-log.md. |
Owner's choices during the pre-M1 design review. Details in docs/specs/2026-09-17-pre-m1-design.md. |
| 2026-09-17 |
Threat model: the main adversary is injected text steering the model. Secondary: any one role process or tool container is compromised, and the goal is containment. brokerd trusts nothing loopd reports beyond the request itself and tracks session taint on its own. |
Owner's choice. It matches the role split the brief already has. |
| 2026-09-17 |
Mattermost and the tailnet are trusted. A Mattermost compromise presumes the whole machine is compromised and is out of scope. Forged approvals are therefore out of scope, and there is no per-grant approval-path field. An accidental secret or PII leak into Mattermost is an accepted risk in v0. |
Owner: Mattermost sits behind a single-user tailnet with ACLs. |
| 2026-09-17 |
Data classes: every session starts at private. A secret result raises it to secret. public is a provenance label and never lowers a session. The grant that authorises a call names the class of its results (default private); tools never label their own output. A grant's allowed data classes are the highest session taint under which it applies. Unattended egress is controlled by grant mode, not by lowering the floor. |
memory/core.md is in every baseline, so no session is really public. Job-scoped grants are left to M5. |
| 2026-09-17 |
IPC: Unix stream sockets, strict JSON bodies in frames with a 4-byte length prefix and a size cap checked before allocation. Unknown fields rejected, tagged enums, no floats, golden-file fixtures as the wire spec. CBOR and protobuf were considered. |
Every role needs a JSON parser anyway (llama-server, Mattermost, model-written tool arguments), so a binary envelope adds a parser without removing one. Both ends share the proto crate. |
| 2026-09-17 |
Each daemon listens on one socket named after itself (infer.sock, broker.sock, loop.sock, gateway.sock). Peer authentication is by directory permissions per pair of roles. Blocking I/O with threads; no async runtime. |
Small audited dependency tree, and fewer ways for the implementing model to go wrong. |
| 2026-09-17 |
Audit and session logs stay JSONL. Type safety comes from the proto record types and strict decoding; the files are never built by string formatting. Writes go through one writer module so the format can change later. |
Volume is one record per tool decision; fsync and inference dominate the cost by orders of magnitude. |
| 2026-09-17 |
Boxmaker uses the shared llama-server router and the shared Ornith instance. No dedicated instance. inferproxy stays. |
The owner runs coding agents against the same router, and a second resident copy of Ornith does not fit beside Laguna. |
| 2026-09-17 |
The harness runs on straylight, the same host as llama-server. |
Owner's choice. One host to secure, and no inference traffic crosses the tailnet. |
| 2026-09-17 |
Container runtime is rootless Podman. |
It is already the standard runtime on straylight and the owner's other hosts. |
| 2026-09-17 |
Implementation is done by Laguna S 2.1, served by straylight, through OpenCode on the owner's development machine; builds are then deployed to straylight. Design, specs, measurement and review are done by a stronger model. Ornith-1.5-35B-A3B remains the model the harness serves. |
The owner wants to test a local model on real implementation work. Plans must be written as small closed tasks with tests specified up front. |
| 2026-09-17 |
M0 is run by the design model, not by Laguna. |
M0 is measurement and interpretation, and its findings bind the design. |