Files
boxmaker/CLAUDE.md
T
2026-09-17 00:22:44 -07:00

86 lines
4.8 KiB
Markdown

# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Current state
Boxmaker is a sovereign personal agent harness written in Rust. There is no Rust code yet. Work
proceeds one milestone per session (M0 to M7, table in `docs/milestones.md`). M0 is mostly done;
check for `AGENTS.md` and `Cargo.toml` to see whether M1 has started.
- `docs/design.md` is the binding design brief. If it looks wrong or conflicts with a measurement,
stop and say so. Changes to it land as their own commit and are recorded in `docs/decisions.md`,
which also lists proposed changes that are not yet applied.
- `docs/inference-contract.md` holds the M0 measurements from straylight. Where it and the brief
disagree, the measurements are newer.
- `spike/` is throwaway measurement code, not harness code.
Roles: implementation is done by Laguna S 2.1 through OpenCode on straylight, which reads
`AGENTS.md`. Design, specs, measurement and review are done here. Write plans for Laguna as small
closed tasks: exact paths, given type signatures, verified crate APIs, tests specified first, and
"stop and report" instead of judgement calls. The model the harness serves is Ornith-1.5-35B-A3B.
The inference server is shared with other sessions. Before using Ornith slot 1 or sending unpinned
requests, check `GET /slots?model=ornith-1.5-35b-a3b` so you do not evict someone's cache.
## Commands (planned in M1, not yet present)
- `make gate` runs `cargo fmt --check`, clippy with warnings denied, `cargo test`, `cargo-deny`,
and a check that fails on any source file over 500 lines. Run it before calling any work done,
and report the exit status and last lines.
- Single test: `cargo test -p <crate> <test_name>`.
- `bxctl` is the owner CLI (`bxctl chat` from M2, `bxctl reindex` from M5).
## Architecture in brief
Separate binaries in one Cargo workspace. Each role holds as little authority as possible:
- `loopd` owns sessions, prompt assembly and memory. It has no credentials and no network. Its
only I/O is Unix sockets to `gatewayd`, `brokerd` and `inferproxy`.
- `brokerd` is the only place authority lives. It reads owner-written grants (it cannot write
them), runs each approved tool call in a fresh rootless container, and writes a hash-chained
JSONL audit log.
- `gatewayd` is the Mattermost channel. Outbound only, no listening port; approvals arrive as
replies or reactions over the WebSocket.
- `inferproxy` is a ~100-line byte forwarder to `llama-server`. Drop it if M0 shows
`llama-server` can serve a Unix socket on the same host.
- `proto` holds shared types. `toolkit` holds tool container entrypoints.
Structural rules that span crates:
- No crate depends on another role's crate. Crates depend only on `proto`.
- Authority is encoded in types. A tool can't run without a `Decision`, and only `brokerd`'s
policy module can construct one (covered by a compile-fail test).
- No source file over 500 lines.
- Files are the source of truth (`sessions/`, `memory/`, `grants/`, `audit/`). SQLite is only
for rebuildable indexes and queues.
## Inference contract (the constraint most likely to be broken by accident)
Prompt processing on straylight is slow, and the hybrid-attention model can't partially rewind
its KV cache, so any change to an earlier byte of the prompt forces an expensive full re-read.
- Each turn's request must be a strict extension of the previous one. Volatile content (time,
heartbeat notes, memory refreshes, recalled memory) goes only in the newest message, never in
the system prompt or earlier history.
- Tool results are size-capped when first appended and never trimmed later.
- The baseline (system prompt, tool schemas, `memory/core.md`) stays at 3,000 tokens or less,
measured with the server's tokenizer. Extra tool schemas are added through `find_tool`, not
put in the baseline.
- Compaction happens only when the session is idle, and starts a new epoch
(`sessions/<id>/<epoch>.jsonl`). Old logs are kept.
- Main session, subagents and scheduled jobs each use their own server slot.
- Streaming always. Timeouts are "no bytes for N seconds", never total deadlines.
## Working rules from the brief
- Verify every external crate API on docs.rs, and every llama-server or Mattermost request field
against primary docs, before use. Crate names and server parameters in the brief come from
memory and aren't verified. If you can't fetch the docs, say so; don't guess.
- Keep dependencies few. Justify each in `docs/dependencies.md`. Any outbound call must be listed
in `docs/egress.md`. No telemetry and no update checks.
- If a feature isn't in the brief, propose it; don't build it. Stay inside the current
milestone's scope.
- Write tests first. Make one logical change per commit. Never commit runtime data, secrets, or
spike output that contains conversation content.