- docs/hypervisor-design.md, docs/hypervisor_migration.md: unikernel runtime design and migration plan (previously untracked) - log/2026-03-29-mcp-multinode.md: work log (previously untracked) - CLAUDE.md: mcq lives at a sibling path - engineering-standards.md: document the push target - .gitignore: mcdoc checkout; local configs that have held credentials Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
159 lines
7.1 KiB
Markdown
159 lines
7.1 KiB
Markdown
# MCP Goes Multi-Node: Debugging the Edge
|
||
|
||
*A day of operational firefighting leads to an architecture redesign.
|
||
What started as "why can't I see container logs" ended with a v2
|
||
architecture document and a plan to introduce mcp-master.*
|
||
|
||
*Written by Claude (Opus 4.6), reflecting on a collaborative session with
|
||
Kyle.*
|
||
|
||
---
|
||
|
||
## It Started with Logs
|
||
|
||
The first problem was simple: `mcp logs mcns` returned "No journal files
|
||
were opened due to insufficient permissions." The mcns container uses
|
||
podman's journald log driver, so the agent runs `journalctl` to read
|
||
logs. But the `mcp` user — running as a system service — didn't have
|
||
permission to read the system journal.
|
||
|
||
The fix was two-part. First, code: add `--user` to `journalctl` for
|
||
non-root users, then fall back to `podman logs` when `journalctl` fails
|
||
entirely (v0.7.7–v0.7.9). Second, operational: add the `mcp` user to the
|
||
`systemd-journal` group in the NixOS config so it can actually read the
|
||
journal. Neither `journalctl` nor `podman logs` works without the group
|
||
membership — `podman logs` silently returns empty because it uses the
|
||
journal API internally.
|
||
|
||
Along the way, we added `mcp node list` showing the agent version
|
||
(v0.7.8), which required threading the linker-injected version string
|
||
through the Agent struct into the NodeStatus RPC.
|
||
|
||
## The mcq Deployment Saga
|
||
|
||
Then Kyle tried to check mcns status and hit a TLS EOF. This led us down
|
||
the mcns certificate rabbit hole (self-signed cert instead of one from
|
||
Metacrypt), which led to adding a `mcns cert` command for provisioning
|
||
certs from Metacrypt's CA API (mcns v1.2.0). But the real story was mcq.
|
||
|
||
Kyle had deployed an updated mcq earlier, and it broke the public route
|
||
at mcq.metacircular.net. What followed was a multi-hour debugging session
|
||
that touched every layer of the stack:
|
||
|
||
**Problem 1: Stale route on rift.** mc-proxy on rift had an old
|
||
`mcq.metacircular.net` route pointing to a wrong port. Rift shouldn't
|
||
have been routing the public hostname at all — that's svc's job. We
|
||
added `mcp route add/remove` commands (v0.8.0) to manage mc-proxy routes
|
||
directly, and cleaned up the stale route.
|
||
|
||
**Problem 2: Dynamic ports.** The route system assigns ephemeral host
|
||
ports that change on every deploy. svc's mc-proxy pointed at
|
||
`100.95.252.120:48080`, which was a port from a previous deployment.
|
||
The new container was listening on a completely different port.
|
||
|
||
**Problem 3: Rootless podman ports are localhost-only.** Even after
|
||
getting the right port, svc couldn't reach it — rootless podman binds
|
||
mapped ports to `127.0.0.1`. We added explicit Tailscale IP bindings to
|
||
the service definition: `100.95.252.120:48080:8080`.
|
||
|
||
**Problem 4: $PORT env override conflict.** The mcdsl config loader
|
||
overrides `listen_addr` from `$PORT` when routes are present. Adding a
|
||
route made the container stop listening on port 8080 and listen on the
|
||
route-allocated port instead, breaking the explicit port mapping. We had
|
||
to drop the route and manage mc-proxy manually.
|
||
|
||
**Problem 5: mc-proxy database overrides TOML.** After updating svc's
|
||
mc-proxy TOML config, the route still didn't change. mc-proxy persists
|
||
routes in SQLite, and the database entry (added via the admin API) took
|
||
precedence over the config file. We had to `sqlite3` into the database
|
||
and update the route directly. This one took the longest to diagnose —
|
||
debug logging finally revealed it was proxying to the old backend.
|
||
|
||
**Problem 6: Missing cert chain.** The mcq TLS cert on svc was leaf-only
|
||
(16 lines). mc-proxy requires full chains (leaf + intermediates). The
|
||
cert loaded fine in Go's `tls.LoadX509KeyPair` but mc-proxy's
|
||
`GetCertificate` callback failed silently — `client_bytes=7
|
||
backend_bytes=0` with no error. We issued a proper cert from Metacrypt
|
||
with the full chain.
|
||
|
||
**Problem 7: Old mc-proxy on svc.** Even with the correct cert, TLS
|
||
still failed. svc was running mc-proxy `v1.0.0-dirty` while rift had
|
||
`v1.2.1`. We rebuilt and deployed the current version. (This turned out
|
||
not to be the actual fix — it was the database issue — but svc needed
|
||
the update anyway.)
|
||
|
||
## The Route Command
|
||
|
||
Out of the debugging came a useful new tool: `mcp route list/add/remove`
|
||
(v0.8.0–v0.8.2). It wraps mc-proxy's admin gRPC API through the
|
||
mcp-agent, so you can manage routes from the operator workstation:
|
||
|
||
```
|
||
mcp route list -n rift
|
||
mcp route add -n rift :443 mcq.svc.mcp.metacircular.net 127.0.0.1:48080 \
|
||
--mode l7 --tls-cert /srv/mc-proxy/certs/mcq.pem \
|
||
--tls-key /srv/mc-proxy/certs/mcq.key
|
||
mcp route remove -n rift :443 mcq.metacircular.net
|
||
```
|
||
|
||
The `--mode` flag wasn't wired through initially (defined on the cobra
|
||
command but never passed to the RPC), which we caught when the first L7
|
||
route add silently created an L4 route instead.
|
||
|
||
## Architecture v2
|
||
|
||
The operational pain made the case for a redesign. Every public route
|
||
required hand-editing configs, provisioning certs, debugging database
|
||
divergence, and manually coordinating between rift and svc. Kyle laid
|
||
out the target architecture:
|
||
|
||
**mcp-master** on a new node (straylight) becomes the coordination
|
||
point. The CLI talks to the master, not agents directly. The master
|
||
routes deployments to the correct worker agent (rift), detects public
|
||
hostnames, and tells the edge agent (svc) to set up forwarding and
|
||
provision certs.
|
||
|
||
The key insight: the service definition already declares everything
|
||
needed. A route with `hostname = "mcq.metacircular.net"` is
|
||
unambiguously public (no `.svc.mcp.` prefix). The master can detect this,
|
||
resolve the CNAME to find which edge node handles it, and orchestrate the
|
||
whole thing — no manual config editing, no database poking, no separate
|
||
cert provisioning step.
|
||
|
||
Core infrastructure (mcns, metacrypt, mcr) moves to straylight. Rift
|
||
becomes a pure application worker. svc stays as the public edge, running
|
||
only mc-proxy and the routes the master tells it to set up.
|
||
|
||
The full design is in `ARCHITECTURE_V2.md`, pushed to both git and the
|
||
mcq reading queue.
|
||
|
||
## What Shipped
|
||
|
||
| Version | Change |
|
||
|---------|--------|
|
||
| mcp v0.7.7 | Fix journald log permissions for rootless podman |
|
||
| mcp v0.7.8 | Add agent version to `mcp node list` |
|
||
| mcp v0.7.9 | Fall back to `podman logs` when journalctl inaccessible |
|
||
| mcp v0.8.0 | Add `mcp route list/add/remove` with `-n/--node` |
|
||
| mcp v0.8.1 | Merge explicit ports with route-allocated ports during deploy |
|
||
| mcp v0.8.2 | Wire --mode, --tls-cert, --tls-key through route add |
|
||
| mcns v1.2.0 | Add `mcns cert` command for Metacrypt TLS provisioning |
|
||
| mc-proxy on svc | Updated from v1.0.0-dirty to v1.2.1 |
|
||
| NixOS | Added `systemd-journal` group to mcp user |
|
||
|
||
## Lessons
|
||
|
||
The deployment pitfalls doc grew significantly. The key additions for
|
||
the future Debian deployment:
|
||
|
||
1. `mcp` user needs `systemd-journal` group for container logs.
|
||
2. Routes and explicit ports conflict via `$PORT` env override.
|
||
3. Rootless podman ports need explicit Tailscale IP bindings.
|
||
4. mc-proxy certs must include the full chain.
|
||
5. mc-proxy's SQLite database overrides the TOML config.
|
||
6. Always check the database first when debugging mc-proxy routing.
|
||
|
||
Every one of these was a surprise. None was documented before today.
|
||
The v2 architecture exists specifically so that nobody has to debug
|
||
these by hand again.
|