M4a spec: catch-up, recording and state writes, as settled while planning
Which channels are caught up (direct channels with allowed users, allowed channels; a new one starts from now), what posts?since returns at v11.11.0 (changed posts, only those in `order`, at most 1,000), that a post is recorded before it is acted on, that a failed state write stops gatewayd, and that "interrupted" is posted on the first connection after a start only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,10 +1,12 @@
|
|||||||
# M4a design: `gatewayd`, conversations over Mattermost
|
# M4a design: `gatewayd`, conversations over Mattermost
|
||||||
|
|
||||||
Status: approved by the owner, 2026-09-23, with two changes folded in (a warning for a file
|
Status: approved by the owner, 2026-09-23, with two changes folded in (a warning for a file
|
||||||
secret; channels shared with other agents, section 7). M4 is split in two (`docs/decisions.md`): M4a is the
|
secret; channels shared with other agents, section 7). Refined while planning, the same day:
|
||||||
conversation path (this document), M4b is approvals over Mattermost (summarised in section 12,
|
section 9 now says which channels are caught up, what `posts?since` returns, when a post is
|
||||||
specified after M4a is built). Where this document and `docs/design.md` disagree, the brief wins;
|
recorded, and that a failed state write stops `gatewayd`. M4 is split in two (`docs/decisions.md`):
|
||||||
the one change it needs, P15 on secrets, is proposed in `docs/decisions.md`.
|
M4a is the conversation path (this document), M4b is approvals over Mattermost (summarised in
|
||||||
|
section 12, specified after M4a is built). Where this document and `docs/design.md` disagree, the
|
||||||
|
brief wins; the one change it needed, P15 on secrets, is applied (`docs/decisions.md`).
|
||||||
|
|
||||||
Facts this rests on, checked on 2026-09-23 against the owner's server and against Mattermost's API
|
Facts this rests on, checked on 2026-09-23 against the owner's server and against Mattermost's API
|
||||||
source and server source at the server's version, v11.11.0:
|
source and server source at the server's version, v11.11.0:
|
||||||
@@ -278,13 +280,25 @@ every change:
|
|||||||
- `in_flight`: turns sent to `loopd` and not yet answered.
|
- `in_flight`: turns sent to `loopd` and not yet answered.
|
||||||
|
|
||||||
A file that exists but cannot be read or parsed stops `gatewayd`, with
|
A file that exists but cannot be read or parsed stops `gatewayd`, with
|
||||||
`see docs/runbook.md#gateway-state-damaged`; a missing file is a first start.
|
`see docs/runbook.md#gateway-state-damaged`; a missing file is a first start. So does a write that
|
||||||
|
fails while running: going on without a record would answer posts twice after the next restart.
|
||||||
|
|
||||||
|
Only posts in tracked channels are recorded: direct messages, and channels in `allow.channels`. A
|
||||||
|
tracked post is recorded as handled **before** it is acted on, so a crash between the two loses
|
||||||
|
the post rather than answering it twice; the `in_flight` entry reports the loss.
|
||||||
|
|
||||||
**Connecting** (at start and after every loss): `GET /users/me`; open the WebSocket; wait for
|
**Connecting** (at start and after every loss): `GET /users/me`; open the WebSocket; wait for
|
||||||
`hello`. Then **catch up**: for every channel in `channels`, and the direct channel with every
|
`hello`. Then **catch up**: for the direct channel with every allowlisted user
|
||||||
allowlisted user (`POST /channels/direct`), and every channel in `allow.channels`, fetch the posts since the last one handled
|
(`POST /channels/direct`) and every channel in `allow.channels`, fetch the posts since the last one
|
||||||
(`GET /channels/{id}/posts?since=…`) and handle them in `create_at` order, before any live event.
|
handled (`GET /channels/{id}/posts?since=…`) and handle them in `create_at` order, before any live
|
||||||
On a first start, `channels` is empty and nothing is caught up: history is not answered.
|
event. A channel with no mark yet (a first start, or a user or channel newly allowed) is marked
|
||||||
|
"now" and not caught up: history is not answered. A channel no longer allowed is not caught up.
|
||||||
|
|
||||||
|
`posts?since` returns the posts **changed** after that time (Mattermost v11.11.0,
|
||||||
|
`SqlPostStore::GetPostsSince`): only the ids in `order` changed; `posts` also holds the roots of
|
||||||
|
their threads, which may be old; edited and deleted posts come back too. Only posts in `order`
|
||||||
|
created after the mark and not deleted are handled. The server takes at most 1,000 changed posts,
|
||||||
|
in no set order; a full answer is logged as "too many posts to catch up; some may be missed".
|
||||||
|
|
||||||
**Losing the connection**: reconnect after 1, 2, 5, 10, then every 30 seconds, printing one line
|
**Losing the connection**: reconnect after 1, 2, 5, 10, then every 30 seconds, printing one line
|
||||||
per attempt. A REST call answered 429 waits until `X-Ratelimit-Reset`; a 5xx is retried twice; a
|
per attempt. A REST call answered 429 waits until `X-Ratelimit-Reset`; a 5xx is retried twice; a
|
||||||
@@ -293,7 +307,8 @@ a wrong token would not help.
|
|||||||
|
|
||||||
**Restarting**: for every entry left in `in_flight`, `gatewayd` posts "interrupted: gatewayd
|
**Restarting**: for every entry left in `in_flight`, `gatewayd` posts "interrupted: gatewayd
|
||||||
restarted before the answer arrived; ask again" in its thread and removes it. `loopd` finished the
|
restarted before the answer arrived; ask again" in its thread and removes it. `loopd` finished the
|
||||||
turn and logged it; the answer is in the session log.
|
turn and logged it; the answer is in the session log. This happens on the first connection after a
|
||||||
|
start only: after a reconnect, the turns in flight are still running.
|
||||||
|
|
||||||
## 10. Runbook
|
## 10. Runbook
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user