Sixteen tasks by Ornith through the plan driver, the review's fixes, and task 16 for its one low
finding. Checked end to end against the owner's Mattermost.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The test swaps the file between the check and the open through a `between` hook. A reference fix
passed it and the gate (769 tests) in the working tree, caught two mutations, and was removed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A status line splits on single spaces only, and every chunk line must end in CRLF. The bounded
`as` casts in http.rs, handshake.rs and proto's sha1.rs become try_from and from.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A script that rewrote the skeleton's comments put `now_ms`'s text on `post`. Task 14's second
session saw the contradiction and deliberated until cut off. Both comments are fixed, each
todo!() comment in tasks 14 and 15 was read beside its signature, and both tasks now say to stop
and quote a comment that does not fit. Lesson T29.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Task 14's session read the crate to learn the APIs and planned `connect` in prose until it was
cut off. `connect`, `Gateway::new` and `catch_up` are now written; the task lists every call with
its signature and makes the copy and the failing test the first actions. Task 15's comment is
rewritten one step per item. Both checked to fail, then pass when filled from their comments.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Task 08's `header` was one todo!() with a dozen branches; Ornith planned it in its head until the
turn ran out (tip T25). `next_message` and `header` are now written as glue over seven small
helpers, and task 09's `poll` over four, each checked to pass when filled from its comments.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each task's tests were run against a reference at its end state; the end states were replayed
from master in order with the gate at each step (650 to 762 tests); each skeleton compiles
against its tests and fails them. The reference is kept off this machine. Lessons T27 (every
wait in a test has a limit) and T28 (mutate the reference before hand-over) come from this work.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Which channels are caught up (direct channels with allowed users, allowed channels; a new one
starts from now), what posts?since returns at v11.11.0 (changed posts, only those in `order`,
at most 1,000), that a post is recorded before it is acted on, that a failed state write stops
gatewayd, and that "interrupted" is posted on the first connection after a start only.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A secret read from a file warns at startup. In allowed channels and group
messages, Boxmaker answers only posts that name it, or replies in its own
threads that name nobody else; one Boxmaker per machine, each its own bot.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
M4 is split into M4a and M4b. The M4a spec rests on facts checked against
the owner's server and Mattermost's source at v11.11.0: the REST and
WebSocket shapes, and that clients will not post a message starting with
'/'. Records the design decisions, proposes P15 (secrets from a systemd
credential, the environment or an owner-only file), and adds the run-time
rows egress.md was missing since M2a and M3b.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Tasks 01 to 17 by Ornith, the review, and the design model's fixes to the
plan. Tool containers run from a Nix-built image named by digest; checked on
straylight.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
14 moves the pipe handling out of container.rs (a pure move, replayed on its
own); 15 starts threads with Builder and bounds output collection with a 2 s
grace period; 16 escapes container errors in the log and fixes two texts; 17
fixes toolkit's thread start, casts and the egress-proxy form. Each checked
against a reference, which is not kept. Tips T24 to T26 from this run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review findings 1 and 2, both plan defects. curl gains --globoff and a
leading --disable; both podman runs gain --pull=never. The given fetch.rs and
the six golden files change with them. Checked on straylight with a rebuilt
image: a glob URL is one request, a missing image fails at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
On straylight, with real containers from deploy/tools-image.nix, every claim
held: no network without a grant, the limits, the file tools, http_fetch's
host checks including a redirect and a tailnet name, and no leftovers. Two
plan defects found there (curl globbing, podman pulling a missing image),
five lower findings.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The third attempt filled five functions, then ran out of room planning all of
run in one turn. run and run_container are now given; spawn, Io::start,
Io::finish and answer are small todo!()s. Checked fillable: 11 of 11 passed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two sessions ended with nothing written, each out of room while planning the
whole file in one turn. The skeleton has the signatures, the constants and
the steps as comments; the task says to fill one function at a time.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Without it the field was never read in task 11, clippy's dead_code failed the
gate, and #[allow] is forbidden; the implementer stopped on the dilemma.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first attempt at task 08 did both and stopped without a commit. Also asks
the implementer to debug inside the repository, since OpenCode refuses /tmp.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Task 04 added dependencies to toolkit and its git add line left out the lock
file, so the driver stopped on an unclean tree. The lock change is folded into
task 04's commit.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A draft spec for the owner's review and 13 offline tasks with their given
tests: shared tool arguments and host rules in proto, the sealed fetch
target (M3a finding 14), the toolkit tools and SOCKS5 egress proxy, and
brokerd's [runner], podman argument lists, runtime and proxy lifecycle. Each
task's tests were run against a reference at that task's end state (560 to
638 tests, clippy clean); the reference is not in the repository. Adds the
runner-unavailable runbook entry and tip T23 (ETXTBSY in script tests).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Models from the owner: Ornith for tasks 01 to 18, Ornith then Grok 4.6 for
19, Grok 4.6 for 20 to 22. Future and late dates set to the commit dates;
notes pasted into the wrong rows moved back; every row has eight cells.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Tasks 01 to 22, the review, and task 23 (the review fixes, with a second,
independent review of them). The one conflict, docs/implementer-lessons.md,
had the same T18 and T19 on both sides; m3a's T20 to T22 follow them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The review table gains findings 16 to 20, which the two review agents
reported and the first write-up dropped. The independent review of the fix
commits, and what was changed for it, is recorded; task 23's claims about
its tests are corrected. The spec and decisions record the day-long cap, the
ttl_ms bound, the socket-directory rule and the listener's retry.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A .jsonl file whose name is not a date is ignored by brokerd and by verify;
it is now listed, so "ok" does not seem to cover it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A plain start refuses, --accept-break records the break and serves, and the
next plain start serves. The given tests covered this only at library level.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From the independent review of task 23. serde quotes a bad frame's text after
decoding, so a compromised peer could put escape sequences in it. A timed-out
admin request now says whether brokerd acted is unknown, since it may have.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
From the independent review of task 23. A torn last line followed by an empty
later file had its recovery written into the later file, which broke the
chain for good; the line is now ended in its own file. The log-name rule
takes months 01 to 12 and days 01 to 31 only. [approvals] ttl_ms is limited
to a day, the longest loopd waits after a pending frame. Running out of file
descriptors or memory pauses the listener instead of stopping brokerd (the
errors the previous fix skipped do not occur on Linux). args.rs's doc fixed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
brokerd makes a socket's directory 0700. With a socket directly in / it would
chmod /, and through a symbolic link it would change the link's target. Both
are now refused at start with #brokerd-start-failed. Without the fix the link
case started and served, with the shared directory made private.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
loopd's self-test caught the change (context per slot 131072 -> 262144,
slots 2 -> 4). The device tests keep the expectation in one constant, and
the M3a script matches it; verify-device passes 6 of 6 and the M3a device
check passes. The inference contract notes which M0 findings rest on the old
layout and need re-measuring.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
M3a review findings 9, 10, 12, 13. A retrying error and every error detail
can carry the inference server's body, so they are escaped like model text.
Admin requests wait at most 30 s, so a stuck brokerd cannot hang bxctl or a
chat turn. A failed write is AdminError::Io and stops handle_pending instead
of being answered with another write. The usage line says what audit verify
checks.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
M3a review finding 8. The wait was expires plus the timeout with no bound,
and the "expiry too far away" guard could not fire, so a far expiry (which
brokerd produces when now + ttl_ms does not fit) parked a turn for ever.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
M3a review finding 7, for the other three roles. loopd keeps its config path
as a path; bxctl and inferproxy take text arguments and answer one that is
not UTF-8 with their usage.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
M3a review findings 3, 5 (the cast), 6, 7 (brokerd), 11. A config, directory
or socket failure at start now ends with docs/runbook.md#brokerd-start-failed,
and losing a listener with #brokerd-listener-lost; both entries are new.
Threads start through thread::Builder, so a refused thread is reported
instead of silently killing a listener; an aborted connection no longer stops
the daemon. brokerd reads args_os and keeps the config path as a path. The
"requester went away" result is recorded at the time it happens.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
M3a review findings 1, 2, 4 and part of 5. One empty log file made brokerd
panic at startup (files[len - 2]); it now opens as an empty log. brokerd's
name check tested one month digit, so a file bxctl ignored could become
brokerd's latest file; both now use proto::is_audit_log_name. Writer no
longer unlinks audit/.lock, which opened a two-writer window. No unwrap in
short_check.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two medium findings in the audit writer (a startup panic on a record-less
log file, and a log-name filter that disagrees with bxctl's), one in the
missing runbook pointers for startup failures, and twelve low ones. Lessons
I14 and T21, T22; two new AGENTS rules; m3a's T18 renumbered to T20 so
master's T18 and T19 survive the merge.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The handoff and the two stopped rows blamed an environment fsync stall.
The cause was macOS refusing socket options after the peer closes, in
loopd and in the brokerd test client. The handoff now opens with the
resolution, the log gets a review note (the stopped rows are kept as
written), and the lessons gain I13 (measure what you blame, and name
the machine) and T18 (the gate must pass on Talos and on the Mac).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
hold_open (d21baa2) kept each connection open for up to two seconds
after its final frame, reading and dropping anything the peer sent, to
hide a test client that set a read timeout after the handler had closed.
On macOS that call fails with EINVAL; the clients now allow for it
(00a85c1, d7009dc), so the handler goes back to closing at once.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The test client set its read timeout before every frame. On macOS that
fails with EINVAL once the handler has sent its final frame and closed,
so every admin test that reads a second frame failed there (40 runs of
40 at 2408e2c). The frame is already buffered, so the client now takes
that one refusal as the peer having closed and reads it. The plan's copy
changes with it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
macOS refuses every socket option with EINVAL once the peer has closed
(XNU sosetoptlock, bsd/kern/uipc_socket.c), even with unread data still
buffered. loopd set a read timeout before each read, so on macOS every
frame or response that arrived just before the peer closed was lost:
BrokerPort reported the broker unavailable with "os error 22", and the
llama client failed the same way. Twelve loopd test binaries failed on
macOS; Linux never refuses, so the gate on Talos did not see it.
Both places now go through socket::set_read_timeout, which on Apple
targets takes that one refusal as success: a socket shut in both
directions returns its data or the end at once and cannot block. A
zero timeout is still an error.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Debug collection commit for the brokerd admin-test hang investigation (M3a
task 19). Contains:
- crates/brokerd/src/broker.rs: hold_open helper (HOLD_OPEN = 2s read-timeout
loop) applied after forbid and after the final send in broker::handle
- crates/brokerd/src/admin.rs: hold_open applied after forbid and after the
final send in admin::handle
- crates/brokerd/src/grants.rs: check_tool_constraints refactor (match guard
instead of nested if)
- docs/M3a/DEBUG-HANDOFF.md: investigation results added (310 runs, zero
hangs reproduced; stalled fsync cannot be fixed without dropping durability)
The implementer log row was committed separately (a6d81d9).
Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
Investigated the brokerd admin test hang described in docs/M3a/DEBUG-HANDOFF.md.
The hold_open fix is already implemented and eliminates the EINVAL. 310 runs
of the admin test binary (4-thread, 16-thread, disk-stressed, via cargo test)
produced zero hangs. The ~7-12% fsync stall cannot be reproduced in this
environment. Per AGENTS.md point 4 for stopped tasks, logging the blocker and
committing only the implementer log.
Implemented-By: OpenCode session (model recorded in docs/implementer-log.md)
The first M3a run's orchestrator copied a reference file from another
checkout and logged it as its own work, and its workers got a summary
of AGENTS.md instead of the file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit a301915551)