Add hypervisor design docs and multinode work log; housekeeping

- docs/hypervisor-design.md, docs/hypervisor_migration.md: unikernel
  runtime design and migration plan (previously untracked)
- log/2026-03-29-mcp-multinode.md: work log (previously untracked)
- CLAUDE.md: mcq lives at a sibling path
- engineering-standards.md: document the push target
- .gitignore: mcdoc checkout; local configs that have held credentials

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
2026-09-20 15:06:18 -07:00
co-authored by Claude Fable 5.1
parent c05ed968f9
commit 868ade3f5a
6 changed files with 1269 additions and 4 deletions
+158
View File
@@ -0,0 +1,158 @@
# MCP Goes Multi-Node: Debugging the Edge
*A day of operational firefighting leads to an architecture redesign.
What started as "why can't I see container logs" ended with a v2
architecture document and a plan to introduce mcp-master.*
*Written by Claude (Opus 4.6), reflecting on a collaborative session with
Kyle.*
---
## It Started with Logs
The first problem was simple: `mcp logs mcns` returned "No journal files
were opened due to insufficient permissions." The mcns container uses
podman's journald log driver, so the agent runs `journalctl` to read
logs. But the `mcp` user — running as a system service — didn't have
permission to read the system journal.
The fix was two-part. First, code: add `--user` to `journalctl` for
non-root users, then fall back to `podman logs` when `journalctl` fails
entirely (v0.7.7v0.7.9). Second, operational: add the `mcp` user to the
`systemd-journal` group in the NixOS config so it can actually read the
journal. Neither `journalctl` nor `podman logs` works without the group
membership — `podman logs` silently returns empty because it uses the
journal API internally.
Along the way, we added `mcp node list` showing the agent version
(v0.7.8), which required threading the linker-injected version string
through the Agent struct into the NodeStatus RPC.
## The mcq Deployment Saga
Then Kyle tried to check mcns status and hit a TLS EOF. This led us down
the mcns certificate rabbit hole (self-signed cert instead of one from
Metacrypt), which led to adding a `mcns cert` command for provisioning
certs from Metacrypt's CA API (mcns v1.2.0). But the real story was mcq.
Kyle had deployed an updated mcq earlier, and it broke the public route
at mcq.metacircular.net. What followed was a multi-hour debugging session
that touched every layer of the stack:
**Problem 1: Stale route on rift.** mc-proxy on rift had an old
`mcq.metacircular.net` route pointing to a wrong port. Rift shouldn't
have been routing the public hostname at all — that's svc's job. We
added `mcp route add/remove` commands (v0.8.0) to manage mc-proxy routes
directly, and cleaned up the stale route.
**Problem 2: Dynamic ports.** The route system assigns ephemeral host
ports that change on every deploy. svc's mc-proxy pointed at
`100.95.252.120:48080`, which was a port from a previous deployment.
The new container was listening on a completely different port.
**Problem 3: Rootless podman ports are localhost-only.** Even after
getting the right port, svc couldn't reach it — rootless podman binds
mapped ports to `127.0.0.1`. We added explicit Tailscale IP bindings to
the service definition: `100.95.252.120:48080:8080`.
**Problem 4: $PORT env override conflict.** The mcdsl config loader
overrides `listen_addr` from `$PORT` when routes are present. Adding a
route made the container stop listening on port 8080 and listen on the
route-allocated port instead, breaking the explicit port mapping. We had
to drop the route and manage mc-proxy manually.
**Problem 5: mc-proxy database overrides TOML.** After updating svc's
mc-proxy TOML config, the route still didn't change. mc-proxy persists
routes in SQLite, and the database entry (added via the admin API) took
precedence over the config file. We had to `sqlite3` into the database
and update the route directly. This one took the longest to diagnose —
debug logging finally revealed it was proxying to the old backend.
**Problem 6: Missing cert chain.** The mcq TLS cert on svc was leaf-only
(16 lines). mc-proxy requires full chains (leaf + intermediates). The
cert loaded fine in Go's `tls.LoadX509KeyPair` but mc-proxy's
`GetCertificate` callback failed silently — `client_bytes=7
backend_bytes=0` with no error. We issued a proper cert from Metacrypt
with the full chain.
**Problem 7: Old mc-proxy on svc.** Even with the correct cert, TLS
still failed. svc was running mc-proxy `v1.0.0-dirty` while rift had
`v1.2.1`. We rebuilt and deployed the current version. (This turned out
not to be the actual fix — it was the database issue — but svc needed
the update anyway.)
## The Route Command
Out of the debugging came a useful new tool: `mcp route list/add/remove`
(v0.8.0v0.8.2). It wraps mc-proxy's admin gRPC API through the
mcp-agent, so you can manage routes from the operator workstation:
```
mcp route list -n rift
mcp route add -n rift :443 mcq.svc.mcp.metacircular.net 127.0.0.1:48080 \
--mode l7 --tls-cert /srv/mc-proxy/certs/mcq.pem \
--tls-key /srv/mc-proxy/certs/mcq.key
mcp route remove -n rift :443 mcq.metacircular.net
```
The `--mode` flag wasn't wired through initially (defined on the cobra
command but never passed to the RPC), which we caught when the first L7
route add silently created an L4 route instead.
## Architecture v2
The operational pain made the case for a redesign. Every public route
required hand-editing configs, provisioning certs, debugging database
divergence, and manually coordinating between rift and svc. Kyle laid
out the target architecture:
**mcp-master** on a new node (straylight) becomes the coordination
point. The CLI talks to the master, not agents directly. The master
routes deployments to the correct worker agent (rift), detects public
hostnames, and tells the edge agent (svc) to set up forwarding and
provision certs.
The key insight: the service definition already declares everything
needed. A route with `hostname = "mcq.metacircular.net"` is
unambiguously public (no `.svc.mcp.` prefix). The master can detect this,
resolve the CNAME to find which edge node handles it, and orchestrate the
whole thing — no manual config editing, no database poking, no separate
cert provisioning step.
Core infrastructure (mcns, metacrypt, mcr) moves to straylight. Rift
becomes a pure application worker. svc stays as the public edge, running
only mc-proxy and the routes the master tells it to set up.
The full design is in `ARCHITECTURE_V2.md`, pushed to both git and the
mcq reading queue.
## What Shipped
| Version | Change |
|---------|--------|
| mcp v0.7.7 | Fix journald log permissions for rootless podman |
| mcp v0.7.8 | Add agent version to `mcp node list` |
| mcp v0.7.9 | Fall back to `podman logs` when journalctl inaccessible |
| mcp v0.8.0 | Add `mcp route list/add/remove` with `-n/--node` |
| mcp v0.8.1 | Merge explicit ports with route-allocated ports during deploy |
| mcp v0.8.2 | Wire --mode, --tls-cert, --tls-key through route add |
| mcns v1.2.0 | Add `mcns cert` command for Metacrypt TLS provisioning |
| mc-proxy on svc | Updated from v1.0.0-dirty to v1.2.1 |
| NixOS | Added `systemd-journal` group to mcp user |
## Lessons
The deployment pitfalls doc grew significantly. The key additions for
the future Debian deployment:
1. `mcp` user needs `systemd-journal` group for container logs.
2. Routes and explicit ports conflict via `$PORT` env override.
3. Rootless podman ports need explicit Tailscale IP bindings.
4. mc-proxy certs must include the full chain.
5. mc-proxy's SQLite database overrides the TOML config.
6. Always check the database first when debugging mc-proxy routing.
Every one of these was a surprise. None was documented before today.
The v2 architecture exists specifically so that nobody has to debug
these by hand again.