- docs/hypervisor-design.md, docs/hypervisor_migration.md: unikernel runtime design and migration plan (previously untracked) - log/2026-03-29-mcp-multinode.md: work log (previously untracked) - CLAUDE.md: mcq lives at a sibling path - engineering-standards.md: document the push target - .gitignore: mcdoc checkout; local configs that have held credentials Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
7.1 KiB
MCP Goes Multi-Node: Debugging the Edge
A day of operational firefighting leads to an architecture redesign. What started as "why can't I see container logs" ended with a v2 architecture document and a plan to introduce mcp-master.
Written by Claude (Opus 4.6), reflecting on a collaborative session with Kyle.
It Started with Logs
The first problem was simple: mcp logs mcns returned "No journal files
were opened due to insufficient permissions." The mcns container uses
podman's journald log driver, so the agent runs journalctl to read
logs. But the mcp user — running as a system service — didn't have
permission to read the system journal.
The fix was two-part. First, code: add --user to journalctl for
non-root users, then fall back to podman logs when journalctl fails
entirely (v0.7.7–v0.7.9). Second, operational: add the mcp user to the
systemd-journal group in the NixOS config so it can actually read the
journal. Neither journalctl nor podman logs works without the group
membership — podman logs silently returns empty because it uses the
journal API internally.
Along the way, we added mcp node list showing the agent version
(v0.7.8), which required threading the linker-injected version string
through the Agent struct into the NodeStatus RPC.
The mcq Deployment Saga
Then Kyle tried to check mcns status and hit a TLS EOF. This led us down
the mcns certificate rabbit hole (self-signed cert instead of one from
Metacrypt), which led to adding a mcns cert command for provisioning
certs from Metacrypt's CA API (mcns v1.2.0). But the real story was mcq.
Kyle had deployed an updated mcq earlier, and it broke the public route at mcq.metacircular.net. What followed was a multi-hour debugging session that touched every layer of the stack:
Problem 1: Stale route on rift. mc-proxy on rift had an old
mcq.metacircular.net route pointing to a wrong port. Rift shouldn't
have been routing the public hostname at all — that's svc's job. We
added mcp route add/remove commands (v0.8.0) to manage mc-proxy routes
directly, and cleaned up the stale route.
Problem 2: Dynamic ports. The route system assigns ephemeral host
ports that change on every deploy. svc's mc-proxy pointed at
100.95.252.120:48080, which was a port from a previous deployment.
The new container was listening on a completely different port.
Problem 3: Rootless podman ports are localhost-only. Even after
getting the right port, svc couldn't reach it — rootless podman binds
mapped ports to 127.0.0.1. We added explicit Tailscale IP bindings to
the service definition: 100.95.252.120:48080:8080.
Problem 4: $PORT env override conflict. The mcdsl config loader
overrides listen_addr from $PORT when routes are present. Adding a
route made the container stop listening on port 8080 and listen on the
route-allocated port instead, breaking the explicit port mapping. We had
to drop the route and manage mc-proxy manually.
Problem 5: mc-proxy database overrides TOML. After updating svc's
mc-proxy TOML config, the route still didn't change. mc-proxy persists
routes in SQLite, and the database entry (added via the admin API) took
precedence over the config file. We had to sqlite3 into the database
and update the route directly. This one took the longest to diagnose —
debug logging finally revealed it was proxying to the old backend.
Problem 6: Missing cert chain. The mcq TLS cert on svc was leaf-only
(16 lines). mc-proxy requires full chains (leaf + intermediates). The
cert loaded fine in Go's tls.LoadX509KeyPair but mc-proxy's
GetCertificate callback failed silently — client_bytes=7 backend_bytes=0 with no error. We issued a proper cert from Metacrypt
with the full chain.
Problem 7: Old mc-proxy on svc. Even with the correct cert, TLS
still failed. svc was running mc-proxy v1.0.0-dirty while rift had
v1.2.1. We rebuilt and deployed the current version. (This turned out
not to be the actual fix — it was the database issue — but svc needed
the update anyway.)
The Route Command
Out of the debugging came a useful new tool: mcp route list/add/remove
(v0.8.0–v0.8.2). It wraps mc-proxy's admin gRPC API through the
mcp-agent, so you can manage routes from the operator workstation:
mcp route list -n rift
mcp route add -n rift :443 mcq.svc.mcp.metacircular.net 127.0.0.1:48080 \
--mode l7 --tls-cert /srv/mc-proxy/certs/mcq.pem \
--tls-key /srv/mc-proxy/certs/mcq.key
mcp route remove -n rift :443 mcq.metacircular.net
The --mode flag wasn't wired through initially (defined on the cobra
command but never passed to the RPC), which we caught when the first L7
route add silently created an L4 route instead.
Architecture v2
The operational pain made the case for a redesign. Every public route required hand-editing configs, provisioning certs, debugging database divergence, and manually coordinating between rift and svc. Kyle laid out the target architecture:
mcp-master on a new node (straylight) becomes the coordination point. The CLI talks to the master, not agents directly. The master routes deployments to the correct worker agent (rift), detects public hostnames, and tells the edge agent (svc) to set up forwarding and provision certs.
The key insight: the service definition already declares everything
needed. A route with hostname = "mcq.metacircular.net" is
unambiguously public (no .svc.mcp. prefix). The master can detect this,
resolve the CNAME to find which edge node handles it, and orchestrate the
whole thing — no manual config editing, no database poking, no separate
cert provisioning step.
Core infrastructure (mcns, metacrypt, mcr) moves to straylight. Rift becomes a pure application worker. svc stays as the public edge, running only mc-proxy and the routes the master tells it to set up.
The full design is in ARCHITECTURE_V2.md, pushed to both git and the
mcq reading queue.
What Shipped
| Version | Change |
|---|---|
| mcp v0.7.7 | Fix journald log permissions for rootless podman |
| mcp v0.7.8 | Add agent version to mcp node list |
| mcp v0.7.9 | Fall back to podman logs when journalctl inaccessible |
| mcp v0.8.0 | Add mcp route list/add/remove with -n/--node |
| mcp v0.8.1 | Merge explicit ports with route-allocated ports during deploy |
| mcp v0.8.2 | Wire --mode, --tls-cert, --tls-key through route add |
| mcns v1.2.0 | Add mcns cert command for Metacrypt TLS provisioning |
| mc-proxy on svc | Updated from v1.0.0-dirty to v1.2.1 |
| NixOS | Added systemd-journal group to mcp user |
Lessons
The deployment pitfalls doc grew significantly. The key additions for the future Debian deployment:
mcpuser needssystemd-journalgroup for container logs.- Routes and explicit ports conflict via
$PORTenv override. - Rootless podman ports need explicit Tailscale IP bindings.
- mc-proxy certs must include the full chain.
- mc-proxy's SQLite database overrides the TOML config.
- Always check the database first when debugging mc-proxy routing.
Every one of these was a surprise. None was documented before today. The v2 architecture exists specifically so that nobody has to debug these by hand again.