docs/design.md is the kickoff pack as written: design brief, session 1 prompt, and milestone outline. Committed unmodified so later edits to the brief show up as diffs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
19 KiB
Boxmaker: kickoff pack
Boxmaker is a sovereign personal agent harness. The name is from Count Zero: the remnant AI in the Straylight cores that patiently assembles small boxes from fragments.
Three parts:
- Design brief — commit as
docs/design.mdin the new repo. Every session reads it. - Session 1 prompt — paste into a fresh Opus session in an empty repo directory.
- Milestone outline and per-milestone prompt template — one session per milestone.
Crate names and llama.cpp server parameters below are from memory or secondhand research. The prompts tell the session to verify each one against primary docs before use.
Part 1: Design brief (docs/design.md)
Goal
A personal agent harness that maximizes sovereignty: local inference by default, no ambient authority, state in plain files the owner can read, a codebase small enough to audit.
Target environment
- Inference host:
straylight, AMD Strix Halo, 128 GB unified memory, llama.cppllama-server, ~128k context configured. - Model: Ornith-1.5-35B-A3B (Qwen3.5-family MoE, hybrid linear/full attention, ~20 KB/token KV at f16, Qwen XML tool calls with thinking blocks).
- Memory is abundant. Generation is fast. Prompt processing is slow and cache invalidation is the dominant cost. Hybrid-attention models cannot partially rewind their cache.
- Reference numbers for the same model family on Strix Halo (Qwen3.6-35B-A3B Q4, llama.cpp, community benchmark grid, May 2026): prompt processing ~1,100 tokens/s at depth 0 and ~700 at 32k depth; generation ~60 tokens/s falling to ~49. M0 replaces these with measurements from straylight. At these rates an uncached 18k-token prompt costs 16–25 s and a 64k re-read costs well over a minute.
- Deployment: rootless containers on a host that is already a tailnet node. Containers do not get their own tailnet identities, so tailnet ACLs cannot tell one role from another. Isolation between roles is done with container networking on the host (see "Network isolation"). Tailnet ACLs restrict what the host as a whole may reach.
- First channel: the owner's existing self-hosted Mattermost server, reached over the tailnet.
Non-goals (v0)
Plugin registry, skill marketplace, web dashboard, multi-user, voice, more than one chat channel, cloud-led orchestration as the default, in-process third-party code of any kind.
Roles (separate binaries, narrow interfaces)
| Role | Holds | Can reach | Never holds |
|---|---|---|---|
gatewayd |
Mattermost bot token, allowlist, message queue | Mattermost server (outbound only), loopd socket |
Tool secrets, grants |
loopd |
Session logs, prompt assembly, memory files | Three Unix sockets: gatewayd, brokerd, inferproxy |
Any credential, any network interface |
inferproxy |
The llama-server address |
llama-server only |
Anything else; it forwards bytes and logs nothing |
brokerd |
Grants (read-only), secrets, audit log | Container runtime, gatewayd socket (for approvals) |
Conversation history |
| tool containers | Only what one call was granted | Only what one call was granted | Anything persistent |
v0 runs all roles on one host. The interfaces must not assume co-location with llama-server.
Network isolation
The host's tailnet identity is shared by every container that has a network, so a container with ordinary outbound networking can reach the whole tailnet and the internet as the host.
loopdruns with no network namespace interfaces except loopback (--network=none). Its only I/O is three Unix sockets on a shared volume.inferproxyis a forwarder of about 100 lines: one Unix socket in, one fixed upstream out. If M0 showsllama-servercan listen on a Unix socket and it runs on the same host, dropinferproxyand mount that socket instead.gatewaydhas outbound network only. It opens no listening TCP port. v0 approvals use replies or emoji reactions, which arrive over the Mattermost WebSocket, so no callback URL is needed.- Tool containers get no network unless the grant names hosts. Granted network access goes through
an allowlisting egress proxy that
brokerdconfigures per call. - Tailnet ACLs (kept in
deploy/) limit the host's node tollama-serverand Mattermost, plus whatever tool grants need. This is defence in depth, not the role boundary.
Channel: Mattermost
- Bot account with a token held in the
SecretStore. REST API for posting, WebSocket for events. Implement directly over HTTP and WebSocket; do not adopt a community SDK without checking its maintenance. Fetch the Mattermost API reference before writing any request. - Identity is the Mattermost user ID. The allowlist is a config file of user IDs. Direct messages only by default; channels must be allowlisted by ID. Authentication is delegated to the server.
- Long tasks: post a placeholder, edit or reply when finished. Threads map to sessions.
- Mattermost stores messages unencrypted in its database. That is acceptable because the server is
the owner's, and it is recorded in
docs/egress.md.
Inference contract (the part existing harnesses get wrong)
- Baseline budget. System prompt plus tool schemas at session start: 3,000 tokens or less, measured with the server's tokenizer, enforced by a test.
- Append-only. The request for turn N+1 is a strict extension of the request for turn N. Nothing volatile (time, heartbeat notes, memory refreshes) goes anywhere but the newest message. Enforced by a property test on the serialized message array.
- No in-place pruning. Tool results are size-capped when first appended, never trimmed later.
- Compaction is an epoch change. It happens only when the session is idle, writes a summary, and starts a new epoch whose prefix is the baseline plus the summary. The old log is kept.
- Pinned slots. Main session, subagents, and scheduled runs each use their own server slot so they do not evict each other's cache. Verify the pinning mechanism on the deployed build.
- Progressive disclosure. Only a small core tool set has schemas in the baseline. Everything
else is found through one
find_toolcall that appends the schema when needed. - Liveness, not deadlines. Streaming always. The timeout is "no bytes for N seconds", never a total-request deadline.
- Runaway control. Per-turn thinking-token cap, per-turn tool-iteration cap, detection of repeated identical tool calls.
- Startup self-test. On boot
loopdchecks: tool-call round trip parses, turn-2 prompt processing count shows a cache hit, configured context matches what the server reports. It refuses to start if any check fails. - Serving settings are part of the repo. Launch flags, chat template hash, sampling settings (temp 0.6, top-p 0.95, top-k 20), f16 KV cache, no speculative decoding.
Open question settled by measurement in M0: use the server's chat-completions endpoint with server-side tool parsing, or render the template in-process for byte-exact prefix control. Default to chat-completions. Switch only if the M0 cache measurements fail. Known risk: Qwen-style templates drop earlier thinking blocks from history, which changes the rendered prefix at the last assistant turn.
Authority contract
loopdhas no authority. Every tool call goes tobrokerdas a request.- A grant is a file the owner writes: tool, argument constraints (paths, hosts, patterns),
allowed data classes, expiry, mode (
auto,ask,deny).brokerdreads grants and cannot write them.loopdcannot see the directory. - No matching grant means deny.
askroutes an approval request to the owner throughgatewaydand the turn suspends until answered or expired. - Each approved call runs in a fresh rootless container: no network unless granted, only granted paths mounted, only the granted secret injected.
- Audit log: append-only JSONL, hash-chained, one record per decision including denials.
- Data classes: every tool result is labelled
public,private, orsecret. A session's taint is the union of labels it has seen. The cloud-consult tool is refused for tainted sessions unless a grant allows that class out. Every outbound cloud payload is logged in full. - Scheduled jobs are re-validated against current grants on every run.
Cloud consult (M6)
- One tool,
consult, behindbrokerdlike any other. The endpoint is any OpenAI-compatible URL, set in config. Nothing in the code is specific to one vendor. - Candidate as of 2026-09-16: GLM-5.3-Flash (open weights, MIT, 320B total / 18B active). It does not run usefully on the inference host, so it is cloud-only here. Hosting choices: Z.ai's international API (Singapore entity, Singapore law, states it does not store API content), or a third-party host chosen by jurisdiction with zero-data-retention routing. Z.ai's subscription coding plan is restricted to its officially supported tools; use pay-as-you-go.
- Call it at temperature 0.6 or lower. Community tests show malformed tool calls at the default
of 1.0. Prefer giving
consultno tools at all: it answers a self-contained question in text. - Hosted GLM always thinks. Treat
consultas slow and asynchronous. - Payloads are built from an explicit allowlist of fields, never from raw session history.
Memory (M5)
Memory is what carries continuity across epoch compaction, so it is required. It must not break the inference contract, and it must not become a way for injected text to persist.
- Core block.
memory/core.md, owner-curated, 300 tokens or less, part of the baseline. Edits take effect at the next epoch, never mid-session. The agent can propose changes; only anaskapproval or the owner's editor applies them. - Everything else is retrieved, not injected. A
recall(query, k)tool returns size-capped snippets with file, line and provenance. Results are appended like any tool result, so the cache prefix is untouched. - Notes.
memory/notes/YYYY-MM-DD.md, append-only, written by aremembertool. Each entry records its session, time, data-class label and whether the session had seen untrusted content. - Flush before compaction. When a session goes idle and is about to change epoch, a job on the scheduled slot extracts durable facts from the log into notes. The main slot is not used.
- Provenance on recall. Recalled text is presented as data with its provenance, never placed in the system prompt. Entries written from sessions that saw untrusted content are marked as such when recalled. Recalled entries carry their data-class label and taint the session.
- Both tools go through
brokerdunder ordinary grants:recalland note-writingauto, core and topic editsask. - Index. Lexical search first: SQLite FTS5 over the Markdown files, rebuildable with
bxctl reindex. Retrieval sits behind aRetrievertrait. - Embeddings are a later, additive index (M5b), never a store. Local model only, served by a
separate
llama-serverembedding instance so chat slots are unaffected. Vectors live in a flat file or SQLite table keyed by chunk hash and embedding-model hash; a model change triggers a rebuild. Combine with lexical results by rank fusion. Add M5b when loggedrecallmisses show lexical search failing, not before. - No vector database service. Search is exact brute-force cosine, in-process. Revisit only if
the index passes about one million chunks or measured
recalllatency exceeds one second, and then prefer an embedded, file-backed approximate index over a server.
State
Files are the source of truth. SQLite is allowed only for rebuildable indexes and queues.
sessions/<id>/<epoch>.jsonl— append-only session logmemory/core.md,memory/notes/*.md,memory/topics/*.md— owner-readable memoryindex/— FTS5 and, later, vectors; disposable and rebuildable frommemory/grants/*.toml— owner-writtenaudit/*.jsonl— hash-chained- Secrets: behind a
SecretStoretrait. v0 backend is an encrypted file whose key is not stored beside it. No plaintext secrets on disk, none in config, none in the repo.
Code constraints
- Rust stable, Cargo workspace. Crates:
proto(shared types),loopd,brokerd,gatewayd,inferproxy,toolkit(tool container entrypoints),bxctl(owner CLI). - No source file over 500 lines. No crate depends on another role's crate, only on
proto. - Authority is encoded in types: a tool cannot execute without a
Decisionvalue that onlybrokerd's policy module can construct. - Dependencies are few and justified in
docs/dependencies.md.cargo-denyruns in the gate. - No telemetry, no update checks, no outbound call not listed in
docs/egress.md.
Part 2: Session 1 prompt (M0 spike + M1 skeleton)
You are starting a new Rust project: Boxmaker, a sovereign personal agent harness. Read docs/design.md
first; it is the design brief and it is binding. If anything in it looks wrong or contradicts
what you measure, stop and tell me rather than working around it.
This session has two deliverables, in order.
DELIVERABLE 1 — M0 measurement spike (throwaway code, kept findings)
The inference host is `straylight` (Strix Halo, llama.cpp llama-server, Ornith-1.5-35B-A3B,
~128k context). Before any architecture is built, measure what the design assumes. Ask me for
the server URL and how you can reach it; if you cannot reach it, stop and say exactly what is
missing.
Fetch the llama-server README for the build running on straylight and the model card before
using any endpoint or parameter. Do not rely on memory for request fields. Then measure and
record in docs/inference-contract.md:
a. Prompt-processing and generation throughput at depth 0 and ~32k (llama-bench or timed
requests; say which).
b. A 3-turn conversation with 4 tool schemas and thinking enabled: tokens processed versus
tokens reused from cache on each turn. Show the raw timing fields.
c. Whether tool calls come back correctly parsed through the chat-completions endpoint with
the model's shipped template. Run 20 trials across 4 tools and report the failure count.
d. Whether requests can be pinned to a slot, and whether a second session on another slot
leaves the first slot's cache intact.
e. What happens to the cache when an earlier assistant turn contained a thinking block.
f. The exact launch flags in use, and the flags you recommend, with reasons.
g. Whether llama-server can listen on a Unix socket, and whether it runs on the same host as
the harness containers. This decides whether `inferproxy` is needed.
Spike code goes in spike/ and is labelled throwaway. Conclude with a recommendation on the open
question in the brief: server-side chat-completions, or in-process template rendering.
Stop after M0 and show me the findings before starting M1.
DELIVERABLE 2 — M1 workspace skeleton
- AGENTS.md at the repo root: project standards drawn from docs/design.md "Code constraints",
plus how to run the gate and how to verify against straylight.
- Cargo workspace with the crates named in the brief. `proto` gets the real shared types:
session log records, tool request and result, data-class labels, grant, decision, audit
record. Other crates are minimal binaries that compile.
- Makefile with `make gate`: cargo fmt --check, clippy with warnings denied, cargo test,
cargo-deny, and a check that fails on any source file over 500 lines.
- docs/dependencies.md and docs/egress.md, initially short and accurate.
- Unit tests for `proto` serialization round trips and for the rule that a `Decision` cannot
be constructed outside the policy module (a compile-fail test is acceptable).
Working rules
- Verify every external crate's current API on docs.rs before using it, and check its
maintenance status and license. If you cannot fetch documentation, say so; do not guess.
- Write a short plan before M1 and execute it with subagents. Tests first.
- Run `make gate` before declaring anything done. Report the exit status and last lines.
- One logical change per commit. Never commit runtime data, secrets, or spike output that
contains conversation content.
- Keep it small. If a feature is not in the brief, do not add it. Propose it instead.
Part 3: Milestones (one session each)
| # | Milestone | Proves | Verified by |
|---|---|---|---|
| M0 | Measurement spike | The cache and tool-call assumptions hold on straylight | Numbers in docs/inference-contract.md |
| M1 | Workspace, proto, gate |
The type-level authority model compiles | make gate |
| M2 | loopd core + bxctl chat |
Append-only loop, baseline ≤ 3k tokens, self-test, runaway control, with one fake in-process tool | Prefix-extension property test; turn-2 cache hit on straylight |
| M3 | brokerd |
Grants, deny by default, hash-chained audit, rootless container runner, four tools (read file, write file, sandboxed shell, HTTP fetch with host allowlist), ask approvals through bxctl |
Tests for deny paths; a container with no network grant cannot reach the network |
| M4 | gatewayd + Mattermost |
Bot connection, user-ID allowlist, threads as sessions, async delivery, approvals by reply or reaction, no listening port | End-to-end direct message on the real server; a non-allowlisted user gets no response |
| M5 | Scheduler + memory | Heartbeat and cron on their own slot in isolated sessions; core block, recall and remember with provenance, FTS5 index, flush-then-compact at idle |
Scheduled run leaves the main slot's cache intact; a fact from epoch 1 is recalled in epoch 2; a note written after reading untrusted content is marked on recall |
| M5b | Embedding index (when justified) | Local embedding server, hybrid retrieval behind the Retriever trait, rebuild on model change |
A paraphrased query finds a note that lexical search misses; no memory text leaves the host |
| M6 | Subagents + cloud consult | Subagents on their slot returning summaries; consult tool gated by data-class grants with full payload log | A tainted session is refused; payload log matches what was sent |
| M7 | Split deployment | Each role in its own rootless container; loopd with no network; socket volume permissions; egress proxy for tools; host ACL policy in deploy/ |
From inside the loopd container, requests to the internet, the tailnet and the host all fail; a tool container without a host grant cannot reach the tailnet |
Per-milestone prompt template
Read AGENTS.md, docs/design.md and docs/inference-contract.md. This session delivers milestone
M<n>: <one-line goal from the table>.
Scope: <bullets>. Out of scope: everything else, including later milestones.
Start by restating the acceptance checks from the milestone table as tests or scripted checks.
Write a plan, show it to me, then execute it with subagents, tests first. Verify external APIs
from their documentation before use. Finish with `make gate` and the on-device verification
named in the table; report exit status, last lines, and what you observed on straylight.
If the design brief turns out to be wrong, stop and propose the change to docs/design.md as its
own commit before building on it.
Decisions still open
- Container runtime. Default proposal: rootless Podman.
- Which host runs the harness containers: straylight itself, or another tailnet node.
- Secret store backend beyond the v0 encrypted file.
- Whether cloud-led sessions are ever allowed, and for which data classes.