Add hypervisor design docs and multinode work log; housekeeping
- docs/hypervisor-design.md, docs/hypervisor_migration.md: unikernel runtime design and migration plan (previously untracked) - log/2026-03-29-mcp-multinode.md: work log (previously untracked) - CLAUDE.md: mcq lives at a sibling path - engineering-standards.md: document the push target - .gitignore: mcdoc checkout; local configs that have held credentials Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
@@ -11,4 +11,12 @@
|
|||||||
/mcns
|
/mcns
|
||||||
/mcp
|
/mcp
|
||||||
/mcdeploy
|
/mcdeploy
|
||||||
|
/mcdoc
|
||||||
|
|
||||||
|
|
||||||
|
# Local service configs and tool settings; these have held credentials.
|
||||||
|
/mcat.toml
|
||||||
|
/mcq.toml
|
||||||
|
/mcr.toml
|
||||||
|
/metacrypt.toml
|
||||||
|
/.claude/
|
||||||
|
|||||||
@@ -17,7 +17,6 @@ Metacircular is a multi-service personal infrastructure platform. This root repo
|
|||||||
| `mcat/` | MCIAS login policy tester — lightweight web app to test and audit login policies | Go |
|
| `mcat/` | MCIAS login policy tester — lightweight web app to test and audit login policies | Go |
|
||||||
| `mcdsl/` | Standard library — shared packages for auth, db, config, HTTP/gRPC servers, CSRF, snapshots | Go |
|
| `mcdsl/` | Standard library — shared packages for auth, db, config, HTTP/gRPC servers, CSRF, snapshots | Go |
|
||||||
| `mcdoc/` | Documentation server — renders markdown from Gitea, serves public docs via mc-proxy | Go |
|
| `mcdoc/` | Documentation server — renders markdown from Gitea, serves public docs via mc-proxy | Go |
|
||||||
| `mcq/` | Document review queue — push docs for review, MCP server for Claude integration | Go |
|
|
||||||
| `mcp/` | Control plane — service deployment, container lifecycle, multi-node fleet management (CLI/agent, master in development) | Go |
|
| `mcp/` | Control plane — service deployment, container lifecycle, multi-node fleet management (CLI/agent, master in development) | Go |
|
||||||
| `mcns/` | Networking service — custom Go DNS server, authoritative for internal zones | Go |
|
| `mcns/` | Networking service — custom Go DNS server, authoritative for internal zones | Go |
|
||||||
| `ca/` | PKI infrastructure and secrets for dev/test (not source code, gitignored) | — |
|
| `ca/` | PKI infrastructure and secrets for dev/test (not source code, gitignored) | — |
|
||||||
@@ -26,7 +25,7 @@ Metacircular is a multi-service personal infrastructure platform. This root repo
|
|||||||
|
|
||||||
Each subproject has its own `CLAUDE.md`, `ARCHITECTURE.md`, `Makefile`, and `go.mod`. When working in a subproject, read its own CLAUDE.md first.
|
Each subproject has its own `CLAUDE.md`, `ARCHITECTURE.md`, `Makefile`, and `go.mod`. When working in a subproject, read its own CLAUDE.md first.
|
||||||
|
|
||||||
Some subprojects (mcat, mcdoc, mcq) may live at sibling paths (e.g., `../mcq/`) rather than as subdirectories, depending on workspace layout.
|
Some subprojects live at sibling paths rather than as subdirectories. For example, `mcq` (document review queue) lives at `../mcq/`. This repository contains only core infrastructure services.
|
||||||
|
|
||||||
## Service Dependencies
|
## Service Dependencies
|
||||||
|
|
||||||
@@ -38,7 +37,6 @@ mcias (standalone — no MCIAS dependency)
|
|||||||
├── mc-proxy (uses MCIAS for admin auth)
|
├── mc-proxy (uses MCIAS for admin auth)
|
||||||
├── mcr (uses MCIAS for auth + policy)
|
├── mcr (uses MCIAS for auth + policy)
|
||||||
├── mcdoc (public, no MCIAS — fetches docs from Gitea)
|
├── mcdoc (public, no MCIAS — fetches docs from Gitea)
|
||||||
├── mcq (uses MCIAS for auth; document review queue)
|
|
||||||
├── mcp (uses MCIAS for auth; orchestrates deployment and lifecycle)
|
├── mcp (uses MCIAS for auth; orchestrates deployment and lifecycle)
|
||||||
├── mcns (uses MCIAS for auth; authoritative DNS for internal zones)
|
├── mcns (uses MCIAS for auth; authoritative DNS for internal zones)
|
||||||
└── mcat (tests MCIAS login policies)
|
└── mcat (tests MCIAS login policies)
|
||||||
@@ -56,6 +54,7 @@ make proto # regenerate gRPC code from .proto files
|
|||||||
make proto-lint # buf lint + buf breaking
|
make proto-lint # buf lint + buf breaking
|
||||||
make devserver # build and run locally against srv/ config
|
make devserver # build and run locally against srv/ config
|
||||||
make docker # build container image
|
make docker # build container image
|
||||||
|
make push # push container image to MCR
|
||||||
make clean # remove binaries
|
make clean # remove binaries
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,244 @@
|
|||||||
|
# Hypervisor-Based Service Isolation -- Design Notes
|
||||||
|
|
||||||
|
> **Status**: Brainstorming / future direction. This document is NOT
|
||||||
|
> an active work item. Agents should ignore this document unless
|
||||||
|
> specifically asked to consider it.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The metacircular platform runs Go services as rootless podman
|
||||||
|
containers orchestrated by MCP. This is a pragmatic execution of ideas
|
||||||
|
originally explored in a series of 2015 papers on security kernels,
|
||||||
|
environment isolation, and unikernels (see References). Those papers
|
||||||
|
describe a richer model than what containers provide: hardware-enforced
|
||||||
|
isolation, mandatory inter-environment communication mediation, and
|
||||||
|
capability-based access control. This document explores bridging the
|
||||||
|
gap by running services as unikernel VMs (specifically Nanos) on the
|
||||||
|
MCP control plane.
|
||||||
|
|
||||||
|
## Motivation: What Containers Don't Give Us
|
||||||
|
|
||||||
|
The current platform has the W7 security kernel's three properties --
|
||||||
|
isolated environments, inter-environment communication (IEC), and
|
||||||
|
access mediation -- but implemented cooperatively rather than enforced:
|
||||||
|
|
||||||
|
| W7 Property | Metacircular Today | Enforcement |
|
||||||
|
|---|---|---|
|
||||||
|
| Isolated environments | Rootless podman (namespaces/cgroups) | OS-cooperative -- shared kernel, escape CVEs exist |
|
||||||
|
| IEC | gRPC/TLS through mc-proxy | Application-cooperative -- services *choose* to route through mc-proxy |
|
||||||
|
| Access mediation | MCIAS tokens + per-service policies | Application-level -- services check tokens voluntarily |
|
||||||
|
|
||||||
|
The topology is right. The enforcement mechanism is weak. Containers
|
||||||
|
share a kernel, and any service could bypass mc-proxy to reach the
|
||||||
|
Tailnet directly.
|
||||||
|
|
||||||
|
## What Unikernels Buy Us
|
||||||
|
|
||||||
|
### Hardware-Enforced Isolation
|
||||||
|
|
||||||
|
Each service runs in its own VM with its own kernel. There is no
|
||||||
|
shared kernel to escape from. The security boundary is the hypervisor
|
||||||
|
(KVM), not Linux namespaces. This is the W7 "isolated environments"
|
||||||
|
model enforced by hardware, not convention.
|
||||||
|
|
||||||
|
### Mandatory IEC
|
||||||
|
|
||||||
|
This is the subtle but powerful part. A container on the Tailnet can
|
||||||
|
talk to anything. A unikernel VM with no direct network interface --
|
||||||
|
only a virtio-net device connected to a host-only bridge that the
|
||||||
|
agent controls -- cannot bypass the mediation layer. If the agent is
|
||||||
|
the only Tailnet citizen on the node and VMs can only reach the
|
||||||
|
agent's bridge, then mc-proxy stops being a routing convenience and
|
||||||
|
becomes the IEC mechanism. Communication between environments is
|
||||||
|
mediated by design, not by trust.
|
||||||
|
|
||||||
|
### Reduced TCB
|
||||||
|
|
||||||
|
Container TCB: Linux kernel + podman runtime + container image (often
|
||||||
|
a full distro). Unikernel TCB: KVM + Nanos runtime + the Go binary.
|
||||||
|
No shell, no package manager, no multi-user, no unnecessary syscalls.
|
||||||
|
|
||||||
|
## The Agent as Security Kernel
|
||||||
|
|
||||||
|
The MCP agent is already structurally positioned to be the W7 security
|
||||||
|
kernel for its node. It manages environment lifecycle, controls the
|
||||||
|
IEC layer (mc-proxy routes), provisions credentials (Metacrypt certs),
|
||||||
|
and reports to a central authority (master).
|
||||||
|
|
||||||
|
With unikernels, this role is formalized. The agent becomes the only
|
||||||
|
entity with host access. Services exist in VMs that can only
|
||||||
|
communicate through agent-controlled channels:
|
||||||
|
|
||||||
|
- **Network access**: virtio-net bridge under agent control; the agent
|
||||||
|
decides what each VM can reach.
|
||||||
|
- **Storage access**: 9p/virtio-fs mounts; the agent controls what
|
||||||
|
each VM sees on disk.
|
||||||
|
- **Credential access**: the agent provisions certs into the VM's
|
||||||
|
filesystem before boot.
|
||||||
|
- **Identity**: the agent attests to the master what image hash is
|
||||||
|
running in each VM.
|
||||||
|
|
||||||
|
## Why This Is Feasible for Metacircular
|
||||||
|
|
||||||
|
Several properties of the existing platform make this tractable:
|
||||||
|
|
||||||
|
- **Go + CGO_ENABLED=0**: Every service already produces a static ELF
|
||||||
|
binary. Nanos needs exactly this. The `ops` tool packages them with
|
||||||
|
minimal friction.
|
||||||
|
|
||||||
|
- **Single-process services**: Each service is one Go binary -- no
|
||||||
|
sidecars, no shell scripts, no multi-process orchestration. That is
|
||||||
|
the unikernel sweet spot.
|
||||||
|
|
||||||
|
- **Single-operator trust domain**: No multi-tenant capability
|
||||||
|
delegation or federated attestation needed. The agent is the
|
||||||
|
security kernel for its node; the master is the coordination point.
|
||||||
|
|
||||||
|
- **mc-proxy already mediates traffic**: The routing mesh is already
|
||||||
|
in place. Making it mandatory (rather than optional) for unikernel
|
||||||
|
VMs is an incremental change, not a new system.
|
||||||
|
|
||||||
|
## Design Sketch
|
||||||
|
|
||||||
|
### Runtime Abstraction
|
||||||
|
|
||||||
|
The agent gains a `Runtime` interface. Podman is one implementation;
|
||||||
|
QEMU/KVM is another. Service definitions gain a `runtime` field:
|
||||||
|
|
||||||
|
```toml
|
||||||
|
name = "mcq"
|
||||||
|
runtime = "unikernel" # or "container" (default)
|
||||||
|
tier = "worker"
|
||||||
|
```
|
||||||
|
|
||||||
|
Both runtimes coexist. Services can be converted incrementally.
|
||||||
|
|
||||||
|
### Networking: Host-Only Bridge
|
||||||
|
|
||||||
|
Each unikernel VM gets a virtio-net device on a host-only bridge. The
|
||||||
|
agent runs on the bridge and controls forwarding. VMs cannot reach the
|
||||||
|
Tailnet directly. All external communication flows through mc-proxy on
|
||||||
|
the host.
|
||||||
|
|
||||||
|
This is structurally similar to how rootless podman already works
|
||||||
|
(container ports are localhost-only, mc-proxy routes to them), but
|
||||||
|
with the enforcement moved from convention to network topology.
|
||||||
|
|
||||||
|
### Storage: 9p Passthrough
|
||||||
|
|
||||||
|
Unikernel VMs mount `/srv/<service>/` via QEMU's `-virtfs` 9p
|
||||||
|
passthrough. Writes go directly to the host filesystem. This makes
|
||||||
|
snapshots work the same way as containers -- the agent tars the host
|
||||||
|
directory.
|
||||||
|
|
||||||
|
### Image Building
|
||||||
|
|
||||||
|
Two options (not mutually exclusive):
|
||||||
|
|
||||||
|
1. **Build on agent**: Agent extracts the ELF binary from the OCI
|
||||||
|
image (pulled from MCR) and runs `ops build` locally.
|
||||||
|
2. **Store unikernel images in MCR**: OCI supports arbitrary media
|
||||||
|
types. Unikernel `.img` files could be stored as OCI artifacts.
|
||||||
|
|
||||||
|
Option 1 is simpler to start with. Option 2 is cleaner long-term.
|
||||||
|
|
||||||
|
### Image Attestation
|
||||||
|
|
||||||
|
Before booting a unikernel, the agent hashes the image and reports it
|
||||||
|
to the master. The master compares against expected hashes from the
|
||||||
|
service definition. This is software attestation -- not TPM-based, but
|
||||||
|
it closes the "is this what I deployed?" question. It is a stepping
|
||||||
|
stone toward measured boot with hardware TPM.
|
||||||
|
|
||||||
|
### Snapshot Constraints
|
||||||
|
|
||||||
|
Unikernels have no shell. The `cli` and `exec:` snapshot methods
|
||||||
|
don't work. Only `grpc` snapshots are viable for unikernel services
|
||||||
|
(the service implements the standard `SnapshotService` RPC). The
|
||||||
|
default snapshot method (tar config/db/certs from the host-side 9p
|
||||||
|
mount) works unchanged since the agent tars the host directory, not
|
||||||
|
the VM filesystem.
|
||||||
|
|
||||||
|
### Debugging
|
||||||
|
|
||||||
|
No `podman exec`, no shell. Debugging relies on:
|
||||||
|
|
||||||
|
- Serial console output from QEMU
|
||||||
|
- gRPC health/status endpoints
|
||||||
|
- Structured logging to a file on the 9p mount
|
||||||
|
- The agent can snapshot and inspect VM state
|
||||||
|
|
||||||
|
This is a real loss of convenience. It is the price of proper
|
||||||
|
isolation -- as noted in the 2015 hypervisor paper, "the nature of
|
||||||
|
debugging means that isolation is broken."
|
||||||
|
|
||||||
|
## Difficulty Assessment
|
||||||
|
|
||||||
|
| Aspect | Difficulty | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| Building unikernel images from Go binaries | Easy | Already static ELF, `ops` handles it |
|
||||||
|
| QEMU lifecycle management in agent | Medium | Replace podman calls with qemu-system calls |
|
||||||
|
| Networking (host-only bridge + mc-proxy) | Medium | Similar to rootless podman model |
|
||||||
|
| Persistent storage via 9p | Medium | Well-supported in QEMU, maps to existing `/srv/` layout |
|
||||||
|
| Snapshots | Medium | `grpc` method works; `cli`/`exec` don't |
|
||||||
|
| Image attestation | Medium-Low | SHA-256 of image before boot |
|
||||||
|
| mc-proxy integration | Low | Just needs a reachable IP:port |
|
||||||
|
| Debugging/observability | Annoying | Loss of exec/shell access |
|
||||||
|
|
||||||
|
The minimum meaningful change is the runtime abstraction + isolated
|
||||||
|
networking together. Running a unikernel with full Tailnet access is
|
||||||
|
just a heavier container with worse debugging. The isolation properties
|
||||||
|
only kick in when the agent mediates all communication.
|
||||||
|
|
||||||
|
## Progression Path
|
||||||
|
|
||||||
|
1. **Runtime abstraction in the agent.** `Runtime` interface with
|
||||||
|
podman and qemu implementations. Service definitions gain a
|
||||||
|
`runtime` field. Both coexist.
|
||||||
|
|
||||||
|
2. **Isolated networking for unikernel VMs.** Host-only bridge per
|
||||||
|
node, agent controls forwarding. mc-proxy becomes the mandatory
|
||||||
|
IEC layer for unikernel services.
|
||||||
|
|
||||||
|
3. **Image attestation.** Agent hashes images before boot, reports to
|
||||||
|
master. Master compares against expected values.
|
||||||
|
|
||||||
|
4. **Capability tokens (longer-term).** MCIAS issues operation-scoped
|
||||||
|
tokens instead of identity tokens. The agent's mediation layer
|
||||||
|
enforces them at the network boundary. This is independent of
|
||||||
|
unikernels but synergizes with mandatory mediation.
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
- **Tailscale integration**: Should unikernel VMs ever be first-class
|
||||||
|
Tailnet citizens (via tsnet compiled into the binary), or should the
|
||||||
|
agent always mediate? Mandatory mediation is more secure but means
|
||||||
|
the agent is on the critical path for all traffic.
|
||||||
|
|
||||||
|
- **Resource limits**: QEMU VMs need explicit memory and CPU
|
||||||
|
allocation. The current container model doesn't declare resource
|
||||||
|
requirements. Unikernels would force this.
|
||||||
|
|
||||||
|
- **Mixed fleet**: During transition, some services run as containers
|
||||||
|
and some as unikernels. mc-proxy routes to both. Does the master
|
||||||
|
need to know the runtime type for placement decisions?
|
||||||
|
|
||||||
|
- **ARM support**: Nanos supports aarch64 but the QEMU/KVM story on
|
||||||
|
Raspberry Pi (no KVM on all models) may limit unikernels to amd64
|
||||||
|
nodes.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- Rees, J. "A Security Kernel Based on the Lambda Calculus" (W7
|
||||||
|
security kernel model -- isolated environments, IEC, access
|
||||||
|
mediation)
|
||||||
|
- "Containers, isolation, and operating systems for network spaces"
|
||||||
|
(2015) -- argues the OS must provide a security kernel; unikernels
|
||||||
|
as viable isolation mechanism
|
||||||
|
- "A hypervisor for the modern age" (2015) -- problem statement for a
|
||||||
|
hypervisor providing proper isolation, IEC, and access mediation
|
||||||
|
with a programmatic administrative interface
|
||||||
|
- "A content-addressable data store with object capabilities" (Nebula,
|
||||||
|
2015) -- capability-based access control model
|
||||||
|
- MCP v2 Architecture (`docs/architecture-v2.md`) -- current platform
|
||||||
|
design this document builds on
|
||||||
@@ -0,0 +1,852 @@
|
|||||||
|
# Unikernel Migration Plan
|
||||||
|
|
||||||
|
> **Status**: Detailed work plan. Not an active work item. Agents
|
||||||
|
> should ignore this document unless specifically asked to consider it.
|
||||||
|
>
|
||||||
|
> **Prerequisite**: MCP v2 phase 6 complete -- master running, agents
|
||||||
|
> on all nodes, edge routing, snapshots, and migration all operational.
|
||||||
|
|
||||||
|
## Starting Point
|
||||||
|
|
||||||
|
The MCP agent already has a `runtime.Runtime` interface
|
||||||
|
(`mcp/internal/runtime/runtime.go`) with methods for Pull, Run, Stop,
|
||||||
|
Remove, Inspect, List, Build, Push, ImageExists, and Logs. The only
|
||||||
|
implementation is `Podman` (`mcp/internal/runtime/podman.go`). The
|
||||||
|
agent struct holds `Runtime runtime.Runtime` and all lifecycle
|
||||||
|
operations (deploy, stop, start, undeploy, status) go through this
|
||||||
|
interface.
|
||||||
|
|
||||||
|
This means the runtime abstraction layer is already in place. The
|
||||||
|
migration is primarily: implement a QEMU/Nanos backend for the
|
||||||
|
existing interface, add isolated networking, and extend service
|
||||||
|
definitions with runtime-specific fields.
|
||||||
|
|
||||||
|
## Terminology
|
||||||
|
|
||||||
|
| Term | Meaning |
|
||||||
|
|------|---------|
|
||||||
|
| **VM** | A QEMU/KVM virtual machine running a Nanos unikernel |
|
||||||
|
| **bridge** | A Linux bridge device (`mcp-br0`) on the host for VM networking |
|
||||||
|
| **TAP** | A TAP device attached to the bridge, one per VM |
|
||||||
|
| **9p mount** | QEMU's `-virtfs` passthrough for host directory access |
|
||||||
|
| **ops** | The Nanos toolchain CLI for building unikernel images |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 1: QEMU Runtime Implementation
|
||||||
|
|
||||||
|
**Goal**: A second `runtime.Runtime` implementation that can start and
|
||||||
|
stop Nanos unikernel VMs with basic networking. No isolation
|
||||||
|
enforcement yet -- VMs get host-forwarded ports like containers do.
|
||||||
|
|
||||||
|
### 1.1 NixOS Host Prerequisites
|
||||||
|
|
||||||
|
Add QEMU/KVM packages to the NixOS configuration on rift and orion.
|
||||||
|
svc (Debian) gets equivalent packages via apt.
|
||||||
|
|
||||||
|
Required on all nodes that will run unikernels:
|
||||||
|
|
||||||
|
- `qemu` (specifically `qemu-system-x86_64`)
|
||||||
|
- `ops` CLI (Nanos toolchain) -- install from GitHub release or build
|
||||||
|
from source
|
||||||
|
- KVM access: the `mcp` user needs `/dev/kvm` access. On NixOS, add
|
||||||
|
the user to the `kvm` group. On Debian, same.
|
||||||
|
- `bridge-utils` or `iproute2` for bridge management (Phase 2)
|
||||||
|
|
||||||
|
Verify KVM works: `qemu-system-x86_64 -enable-kvm -nographic
|
||||||
|
-no-reboot` should boot and exit.
|
||||||
|
|
||||||
|
**Deliverable**: All amd64 nodes can run QEMU with KVM acceleration.
|
||||||
|
RPi nodes (arm64, no KVM) are excluded from unikernel support.
|
||||||
|
|
||||||
|
### 1.2 Image Building Pipeline
|
||||||
|
|
||||||
|
The agent needs to produce a Nanos `.img` file from a Go binary. Two
|
||||||
|
paths, implemented in order:
|
||||||
|
|
||||||
|
**1.2a -- Local build from OCI image (initial approach)**
|
||||||
|
|
||||||
|
The agent already pulls OCI images via `Runtime.Pull()`. For
|
||||||
|
unikernels:
|
||||||
|
|
||||||
|
1. Pull the OCI image from MCR (reuse existing podman pull or use
|
||||||
|
`skopeo copy` to a local directory).
|
||||||
|
2. Extract the ELF binary from the image. Convention: the binary is at
|
||||||
|
`/usr/local/bin/<service>` in the image (same path the Dockerfiles
|
||||||
|
use).
|
||||||
|
3. Run `ops build <binary> -c <config.json>` to produce a `.img` file.
|
||||||
|
4. Store the image at `/srv/mcp/images/<service>-<component>.img`.
|
||||||
|
|
||||||
|
The `ops` config JSON specifies:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"Args": ["server", "--config", "/srv/mcq/mcq.toml"],
|
||||||
|
"Dirs": ["srv"],
|
||||||
|
"Mounts": {
|
||||||
|
"/srv/<service>": "/srv/<service>"
|
||||||
|
},
|
||||||
|
"ManifestPassthrough": {
|
||||||
|
"mem": "256m",
|
||||||
|
"smp": 1
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**1.2b -- Pre-built unikernel images in MCR (later)**
|
||||||
|
|
||||||
|
Store `.img` files as OCI artifacts in MCR with a distinct media type
|
||||||
|
(`application/vnd.metacircular.unikernel.nanos.v1`). The agent pulls
|
||||||
|
the artifact and writes it directly to
|
||||||
|
`/srv/mcp/images/<service>-<component>.img`. This skips the
|
||||||
|
extract-and-build step and ensures the deployed image is identical to
|
||||||
|
what was built.
|
||||||
|
|
||||||
|
MCR already stores OCI artifacts; this requires adding the media type
|
||||||
|
to MCR's accepted list and adding an `mcp build --unikernel` command
|
||||||
|
that builds the image locally and pushes it.
|
||||||
|
|
||||||
|
**Deliverable**: Agent can produce a bootable Nanos image from an
|
||||||
|
existing OCI container image.
|
||||||
|
|
||||||
|
### 1.3 QEMU Runtime Type
|
||||||
|
|
||||||
|
Implement `QEMURuntime` satisfying `runtime.Runtime`:
|
||||||
|
|
||||||
|
```go
|
||||||
|
type QEMURuntime struct {
|
||||||
|
imageDir string // /srv/mcp/images/
|
||||||
|
stateDir string // /srv/mcp/vm-state/
|
||||||
|
opsPath string // path to ops binary
|
||||||
|
qemuPath string // path to qemu-system-x86_64
|
||||||
|
logger *slog.Logger
|
||||||
|
mu sync.Mutex
|
||||||
|
vms map[string]*vmState // name → running VM state
|
||||||
|
}
|
||||||
|
|
||||||
|
type vmState struct {
|
||||||
|
pid int
|
||||||
|
qmpSocket string // QMP control socket
|
||||||
|
serial string // serial console log path
|
||||||
|
ip string // VM IP on bridge (Phase 2)
|
||||||
|
ports map[int]int // guest port → host port
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Method mapping:**
|
||||||
|
|
||||||
|
| Runtime Method | QEMU Implementation |
|
||||||
|
|---|---|
|
||||||
|
| `Pull(image)` | Pull OCI image, extract ELF, run `ops build`, store `.img` |
|
||||||
|
| `Run(spec)` | Start `qemu-system-x86_64` with KVM, virtio-net, 9p mounts, QMP socket |
|
||||||
|
| `Stop(name)` | Send `system_powerdown` via QMP, wait 10s, then SIGKILL |
|
||||||
|
| `Remove(name)` | Kill process if running, remove state files |
|
||||||
|
| `Inspect(name)` | Check process liveness + read QMP status |
|
||||||
|
| `List()` | Enumerate `/srv/mcp/vm-state/*/qemu.pid`, check liveness |
|
||||||
|
| `Build(...)` | Not applicable for unikernels (image built during Pull) |
|
||||||
|
| `Push(...)` | Not applicable (future: push `.img` to MCR as OCI artifact) |
|
||||||
|
| `ImageExists(image)` | Check if `.img` file exists in imageDir |
|
||||||
|
| `Logs(name)` | Read serial console log file |
|
||||||
|
|
||||||
|
**QEMU invocation** (Phase 1 -- user-mode networking with port
|
||||||
|
forwards, no bridge yet):
|
||||||
|
|
||||||
|
```
|
||||||
|
qemu-system-x86_64 \
|
||||||
|
-enable-kvm \
|
||||||
|
-m 256 \
|
||||||
|
-smp 1 \
|
||||||
|
-nographic \
|
||||||
|
-serial file:/srv/mcp/vm-state/<name>/console.log \
|
||||||
|
-qmp unix:/srv/mcp/vm-state/<name>/qmp.sock,server,nowait \
|
||||||
|
-drive file=/srv/mcp/images/<name>.img,format=raw,if=virtio \
|
||||||
|
-virtfs local,path=/srv/<service>,mount_tag=srvdata,security_model=mapped-xattr,id=srvdata \
|
||||||
|
-device virtio-net-pci,netdev=net0 \
|
||||||
|
-netdev user,id=net0,hostfwd=tcp:127.0.0.1:<host_port>-:<guest_port>
|
||||||
|
```
|
||||||
|
|
||||||
|
This gives user-mode networking with port forwards to localhost --
|
||||||
|
functionally identical to how rootless podman works. mc-proxy routes
|
||||||
|
to `127.0.0.1:<host_port>` the same way it does for containers.
|
||||||
|
|
||||||
|
**Deliverable**: A `QEMURuntime` that passes the same interface as
|
||||||
|
`Podman`. Agent can deploy, stop, inspect, and undeploy unikernel
|
||||||
|
services using QEMU user-mode networking.
|
||||||
|
|
||||||
|
### 1.4 Service Definition Changes
|
||||||
|
|
||||||
|
Add `runtime` field to service definitions, proto specs, and registry
|
||||||
|
schema.
|
||||||
|
|
||||||
|
**TOML** (`servicedef.go`):
|
||||||
|
|
||||||
|
```toml
|
||||||
|
name = "mcq"
|
||||||
|
runtime = "unikernel"
|
||||||
|
tier = "worker"
|
||||||
|
active = true
|
||||||
|
|
||||||
|
[[components]]
|
||||||
|
name = "mcq"
|
||||||
|
image = "mcr.svc.mcp.metacircular.net:8443/mcq:v0.4.0"
|
||||||
|
memory = 256 # MB, required for unikernels
|
||||||
|
vcpus = 1 # default 1
|
||||||
|
volumes = ["/srv/mcq:/srv/mcq"]
|
||||||
|
cmd = ["server", "--config", "/srv/mcq/mcq.toml"]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Proto** (`mcp.proto`):
|
||||||
|
|
||||||
|
```protobuf
|
||||||
|
message ServiceSpec {
|
||||||
|
string name = 1;
|
||||||
|
bool active = 2;
|
||||||
|
repeated ComponentSpec components = 3;
|
||||||
|
string comment = 4;
|
||||||
|
string runtime = 5; // "container" (default) or "unikernel"
|
||||||
|
}
|
||||||
|
|
||||||
|
message ComponentSpec {
|
||||||
|
// ... existing fields ...
|
||||||
|
int32 memory_mb = 11; // required for unikernel runtime
|
||||||
|
int32 vcpus = 12; // default 1
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Registry schema** (new migration):
|
||||||
|
|
||||||
|
```sql
|
||||||
|
ALTER TABLE components ADD COLUMN runtime TEXT NOT NULL DEFAULT 'container';
|
||||||
|
ALTER TABLE components ADD COLUMN memory_mb INTEGER NOT NULL DEFAULT 0;
|
||||||
|
ALTER TABLE components ADD COLUMN vcpus INTEGER NOT NULL DEFAULT 1;
|
||||||
|
```
|
||||||
|
|
||||||
|
**Agent runtime selection**: In `agent.go`, the agent holds both
|
||||||
|
runtimes:
|
||||||
|
|
||||||
|
```go
|
||||||
|
type Agent struct {
|
||||||
|
// ... existing fields ...
|
||||||
|
ContainerRuntime runtime.Runtime // podman
|
||||||
|
UnikernelRuntime runtime.Runtime // qemu (nil if not configured)
|
||||||
|
}
|
||||||
|
|
||||||
|
func (a *Agent) runtimeFor(comp *registry.Component) runtime.Runtime {
|
||||||
|
if comp.Runtime == "unikernel" {
|
||||||
|
return a.UnikernelRuntime
|
||||||
|
}
|
||||||
|
return a.ContainerRuntime
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
All lifecycle operations call `a.runtimeFor(comp)` instead of
|
||||||
|
`a.Runtime` directly.
|
||||||
|
|
||||||
|
**Validation rules**:
|
||||||
|
- `runtime = "unikernel"` requires `memory_mb > 0`.
|
||||||
|
- `runtime = "unikernel"` requires the node to have KVM
|
||||||
|
(`/dev/kvm` exists). Agent rejects deploys on nodes without KVM.
|
||||||
|
- `runtime = "unikernel"` is incompatible with `exec:` and `cli`
|
||||||
|
snapshot methods. Validation rejects these combinations.
|
||||||
|
|
||||||
|
**Deliverable**: Service definitions can declare `runtime =
|
||||||
|
"unikernel"`. The agent selects the correct runtime per component.
|
||||||
|
Container services are completely unaffected.
|
||||||
|
|
||||||
|
### 1.5 Resource Tracking
|
||||||
|
|
||||||
|
The agent needs to track allocated VM resources to avoid overcommit.
|
||||||
|
The v2 heartbeat already reports CPU, memory, and disk. Add tracking
|
||||||
|
of allocated-to-VMs resources:
|
||||||
|
|
||||||
|
```go
|
||||||
|
type ResourceTracker struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
totalMemMB int64 // from /proc/meminfo
|
||||||
|
totalCPUs int32 // from runtime.NumCPU()
|
||||||
|
allocMemMB int64 // sum of running VM memory_mb
|
||||||
|
allocCPUs int32 // sum of running VM vcpus
|
||||||
|
}
|
||||||
|
|
||||||
|
func (r *ResourceTracker) CanFit(memMB int64, vcpus int32) bool
|
||||||
|
func (r *ResourceTracker) Allocate(memMB int64, vcpus int32)
|
||||||
|
func (r *ResourceTracker) Release(memMB int64, vcpus int32)
|
||||||
|
```
|
||||||
|
|
||||||
|
The master's placement algorithm gains a resource check: before
|
||||||
|
placing a unikernel service on a node, verify the node has enough
|
||||||
|
unallocated memory and CPUs. Container services continue to use
|
||||||
|
container-count placement.
|
||||||
|
|
||||||
|
**Deliverable**: Agent tracks VM resource allocation. Master rejects
|
||||||
|
placements that would overcommit a node.
|
||||||
|
|
||||||
|
### 1.6 Phase 1 Validation
|
||||||
|
|
||||||
|
Deploy a test service (a minimal Go HTTP server, not a real platform
|
||||||
|
service) as a unikernel:
|
||||||
|
|
||||||
|
1. Build a trivial Go binary that serves HTTP on port 8080.
|
||||||
|
2. Package it as an OCI image, push to MCR.
|
||||||
|
3. Write a service definition with `runtime = "unikernel"`.
|
||||||
|
4. `mcp deploy test-unikernel` -- verify it starts, mc-proxy routes
|
||||||
|
to it, health checks pass.
|
||||||
|
5. `mcp undeploy test-unikernel` -- verify clean shutdown.
|
||||||
|
6. Verify container services are completely unaffected.
|
||||||
|
|
||||||
|
**Phase 1 complete when**: A unikernel service can be deployed,
|
||||||
|
health-checked, and undeployed through the normal `mcp deploy`/
|
||||||
|
`mcp undeploy` flow, alongside running container services.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 2: Isolated Networking
|
||||||
|
|
||||||
|
**Goal**: Replace QEMU user-mode networking with a host-only bridge.
|
||||||
|
VMs can only communicate through mc-proxy. This is the phase that
|
||||||
|
delivers the security properties -- without it, unikernels are just
|
||||||
|
heavier containers.
|
||||||
|
|
||||||
|
### 2.1 Bridge Setup
|
||||||
|
|
||||||
|
Create a persistent Linux bridge on each unikernel-capable node:
|
||||||
|
|
||||||
|
**NixOS** (`networking.bridges` in NixOS config):
|
||||||
|
|
||||||
|
```nix
|
||||||
|
networking.bridges.mcp-br0.interfaces = [];
|
||||||
|
networking.interfaces.mcp-br0.ipv4.addresses = [{
|
||||||
|
address = "10.99.0.1";
|
||||||
|
prefixLength = 24;
|
||||||
|
}];
|
||||||
|
```
|
||||||
|
|
||||||
|
**Debian** (svc -- if svc ever runs unikernels, which is unlikely
|
||||||
|
given its edge role, but document for completeness):
|
||||||
|
|
||||||
|
```
|
||||||
|
# /etc/network/interfaces.d/mcp-br0
|
||||||
|
auto mcp-br0
|
||||||
|
iface mcp-br0 inet static
|
||||||
|
address 10.99.0.1/24
|
||||||
|
bridge_ports none
|
||||||
|
bridge_stp off
|
||||||
|
```
|
||||||
|
|
||||||
|
The bridge uses the `10.99.0.0/24` subnet. This is a host-only
|
||||||
|
network -- no default route, no NAT to the internet or Tailnet. VMs
|
||||||
|
can only reach `10.99.0.1` (the agent/mc-proxy host).
|
||||||
|
|
||||||
|
**Deliverable**: Each unikernel-capable node has a `mcp-br0` bridge
|
||||||
|
with address `10.99.0.1/24`.
|
||||||
|
|
||||||
|
### 2.2 TAP Device Management
|
||||||
|
|
||||||
|
Each VM gets a TAP device attached to the bridge. The agent creates
|
||||||
|
and destroys TAP devices as part of the VM lifecycle:
|
||||||
|
|
||||||
|
```go
|
||||||
|
func (q *QEMURuntime) createTAP(name string) (string, error) {
|
||||||
|
tapName := fmt.Sprintf("tap-%s", name) // max 15 chars for IFNAMSIZ
|
||||||
|
// ip tuntap add dev <tap> mode tap user mcp
|
||||||
|
// ip link set <tap> master mcp-br0
|
||||||
|
// ip link set <tap> up
|
||||||
|
return tapName, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
func (q *QEMURuntime) destroyTAP(name string) error {
|
||||||
|
tapName := fmt.Sprintf("tap-%s", name)
|
||||||
|
// ip link del <tap>
|
||||||
|
return nil
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
TAP creation requires `CAP_NET_ADMIN` or `ip tuntap` permissions for
|
||||||
|
the `mcp` user. On NixOS, grant this via a udev rule or by running
|
||||||
|
the agent with ambient capabilities:
|
||||||
|
|
||||||
|
```nix
|
||||||
|
systemd.services.mcp-agent.serviceConfig.AmbientCapabilities = [
|
||||||
|
"CAP_NET_ADMIN"
|
||||||
|
];
|
||||||
|
```
|
||||||
|
|
||||||
|
**QEMU invocation changes** (bridge networking replaces user-mode):
|
||||||
|
|
||||||
|
```
|
||||||
|
qemu-system-x86_64 \
|
||||||
|
... \
|
||||||
|
-device virtio-net-pci,netdev=net0,mac=52:54:00:xx:xx:xx \
|
||||||
|
-netdev tap,id=net0,ifname=tap-<name>,script=no,downscript=no
|
||||||
|
```
|
||||||
|
|
||||||
|
Each VM gets a deterministic MAC address derived from the service
|
||||||
|
name (e.g., SHA-256 of service name, take 5 bytes, prepend `52:54:00`).
|
||||||
|
|
||||||
|
### 2.3 VM IP Assignment
|
||||||
|
|
||||||
|
VMs need static IPs on the bridge. No DHCP server -- the agent
|
||||||
|
assigns IPs and passes them to Nanos via the ops config.
|
||||||
|
|
||||||
|
```go
|
||||||
|
type IPAllocator struct {
|
||||||
|
mu sync.Mutex
|
||||||
|
subnet net.IPNet // 10.99.0.0/24
|
||||||
|
gateway net.IP // 10.99.0.1
|
||||||
|
assigned map[string]net.IP // service name → IP
|
||||||
|
next byte // next octet to try (2-254)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The ops config passes networking to Nanos:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"RunConfig": {
|
||||||
|
"IPAddress": "10.99.0.5",
|
||||||
|
"NetMask": "255.255.255.0",
|
||||||
|
"Gateway": "10.99.0.1"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Assigned IPs are persisted in the agent's registry:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
ALTER TABLE components ADD COLUMN vm_ip TEXT;
|
||||||
|
```
|
||||||
|
|
||||||
|
**Deliverable**: Each VM gets a static IP on the bridge. The agent
|
||||||
|
tracks assignments in its registry.
|
||||||
|
|
||||||
|
### 2.4 mc-proxy Route Update
|
||||||
|
|
||||||
|
With bridge networking, mc-proxy routes change from
|
||||||
|
`127.0.0.1:<host_port>` to `10.99.0.<n>:<guest_port>`:
|
||||||
|
|
||||||
|
- L7 routes: mc-proxy terminates TLS, forwards to
|
||||||
|
`10.99.0.<n>:<port>` (plaintext on the bridge).
|
||||||
|
- L4 routes: mc-proxy passes through to `10.99.0.<n>:<port>` (TLS
|
||||||
|
end-to-end).
|
||||||
|
|
||||||
|
The `ProxyRouter.RegisterRoutes()` method needs to use the VM's
|
||||||
|
bridge IP instead of `127.0.0.1` for unikernel components. Port
|
||||||
|
allocation changes: unikernel VMs expose their actual service port
|
||||||
|
on the bridge (no random host port needed), so `host_port` equals
|
||||||
|
the route's declared port.
|
||||||
|
|
||||||
|
### 2.5 Firewall Rules
|
||||||
|
|
||||||
|
The bridge must be locked down so VMs can only reach mc-proxy:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Allow established connections back to VMs
|
||||||
|
iptables -A FORWARD -i mcp-br0 -o mcp-br0 -m state --state ESTABLISHED,RELATED -j ACCEPT
|
||||||
|
|
||||||
|
# Allow VMs to reach the host (mc-proxy) on the bridge IP
|
||||||
|
iptables -A INPUT -i mcp-br0 -d 10.99.0.1 -j ACCEPT
|
||||||
|
|
||||||
|
# Block VM-to-VM traffic on the bridge
|
||||||
|
ebtables -A FORWARD -i tap-+ -o tap-+ -j DROP
|
||||||
|
|
||||||
|
# Block VMs from reaching anything outside the bridge
|
||||||
|
iptables -A FORWARD -i mcp-br0 ! -o mcp-br0 -j DROP
|
||||||
|
```
|
||||||
|
|
||||||
|
These rules enforce mandatory mediation: VMs can reach the host
|
||||||
|
(where mc-proxy listens) but nothing else. No Tailnet, no internet,
|
||||||
|
no other VMs. All inter-service communication goes through mc-proxy.
|
||||||
|
|
||||||
|
On NixOS, these rules go in `networking.firewall` or
|
||||||
|
`networking.nftables`. On Debian, `/etc/iptables/rules.v4`.
|
||||||
|
|
||||||
|
**Deliverable**: VMs are network-isolated. They can only reach
|
||||||
|
mc-proxy on the host. VM-to-VM and VM-to-Tailnet traffic is blocked.
|
||||||
|
|
||||||
|
### 2.6 Phase 2 Validation
|
||||||
|
|
||||||
|
1. Deploy the test unikernel from Phase 1 with bridge networking.
|
||||||
|
2. Verify mc-proxy routes to it via the bridge IP.
|
||||||
|
3. From inside the VM (via the service's own gRPC or HTTP endpoint),
|
||||||
|
attempt to reach a Tailnet IP directly -- must fail.
|
||||||
|
4. Attempt to reach another VM on the bridge -- must fail.
|
||||||
|
5. Verify the service can reach its dependencies (MCIAS, Metacrypt)
|
||||||
|
only via mc-proxy on the host.
|
||||||
|
6. Verify container services are completely unaffected by the bridge.
|
||||||
|
|
||||||
|
**Phase 2 complete when**: Unikernel VMs are fully network-isolated
|
||||||
|
and can only communicate through mc-proxy. The agent enforces this
|
||||||
|
structurally, not cooperatively.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 3: Snapshots and Observability
|
||||||
|
|
||||||
|
**Goal**: Ensure unikernel services participate in the snapshot and
|
||||||
|
monitoring systems. Adapt debugging tools for the no-shell environment.
|
||||||
|
|
||||||
|
### 3.1 Snapshot Adaptation
|
||||||
|
|
||||||
|
The default snapshot method (tar `*.toml`, `*.db`, `*.pem` from the
|
||||||
|
host-side `/srv/<service>/`) works unchanged for unikernels because
|
||||||
|
the agent tars the host directory, not the VM filesystem. The 9p
|
||||||
|
passthrough means writes from the VM appear on the host immediately.
|
||||||
|
|
||||||
|
The `grpc` snapshot method also works unchanged -- the agent calls the
|
||||||
|
service's `SnapshotService.Snapshot` RPC over mc-proxy, which reaches
|
||||||
|
the VM the same way any other gRPC call does.
|
||||||
|
|
||||||
|
**What doesn't work**: `cli` and `exec:` methods, because there is
|
||||||
|
no shell inside the VM. Validation (from Phase 1.4) already rejects
|
||||||
|
these combinations, but the snapshot scheduler should also log a
|
||||||
|
warning if it encounters a unikernel service with an incompatible
|
||||||
|
snapshot method.
|
||||||
|
|
||||||
|
**Deliverable**: Snapshots work for unikernel services using the
|
||||||
|
default or `grpc` methods.
|
||||||
|
|
||||||
|
### 3.2 Serial Console Log Collection
|
||||||
|
|
||||||
|
QEMU writes serial console output to
|
||||||
|
`/srv/mcp/vm-state/<name>/console.log`. The `Logs()` method on
|
||||||
|
`QEMURuntime` reads this file. But the agent's `Logs` gRPC RPC
|
||||||
|
currently streams from podman/journalctl.
|
||||||
|
|
||||||
|
Extend the `Logs` RPC to detect the component's runtime and read from
|
||||||
|
the serial console log instead:
|
||||||
|
|
||||||
|
```go
|
||||||
|
func (a *Agent) Logs(req *pb.LogsRequest, stream pb.McpAgent_LogsServer) error {
|
||||||
|
comp := a.registryComponent(req)
|
||||||
|
if comp.Runtime == "unikernel" {
|
||||||
|
return a.streamSerialLog(comp, req, stream)
|
||||||
|
}
|
||||||
|
return a.streamContainerLog(comp, req, stream)
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
For Nanos, configure the Go binary's logging to write to stdout/stderr
|
||||||
|
(which Nanos routes to the serial console). This is the default Go
|
||||||
|
behavior, so no changes needed in the services themselves.
|
||||||
|
|
||||||
|
**Deliverable**: `mcp logs <service>` works for unikernel services,
|
||||||
|
streaming the serial console output.
|
||||||
|
|
||||||
|
### 3.3 Health Check Adaptation
|
||||||
|
|
||||||
|
The v2 health check types (tcp, grpc, http) all work over the network
|
||||||
|
and don't require shell access. No changes needed -- the agent's
|
||||||
|
monitoring loop connects to the VM's port via mc-proxy or the bridge
|
||||||
|
IP the same way it does for containers.
|
||||||
|
|
||||||
|
### 3.4 Drift Detection
|
||||||
|
|
||||||
|
The agent's `LiveCheck()` currently calls `Runtime.List()` and
|
||||||
|
reconciles with the registry. The QEMU `List()` implementation
|
||||||
|
enumerates running VMs by checking PIDs in
|
||||||
|
`/srv/mcp/vm-state/*/qemu.pid`. This needs to be reliable:
|
||||||
|
|
||||||
|
- On agent restart, rebuild the `vms` map from the state directory.
|
||||||
|
- QEMU processes started with `--daemonize` survive agent restarts.
|
||||||
|
- The QMP socket reconnects on agent restart.
|
||||||
|
|
||||||
|
**Deliverable**: Drift detection works for unikernel VMs. Agent
|
||||||
|
restart does not lose track of running VMs.
|
||||||
|
|
||||||
|
### 3.5 Phase 3 Validation
|
||||||
|
|
||||||
|
1. Deploy a unikernel service with `[snapshot] method = "grpc"`.
|
||||||
|
2. `mcp snapshot <service>` -- verify snapshot succeeds.
|
||||||
|
3. Verify scheduled snapshots include the unikernel service.
|
||||||
|
4. `mcp logs <service>` -- verify serial console output streams.
|
||||||
|
5. Kill the QEMU process manually. Verify drift detection catches it
|
||||||
|
and reports the service as unhealthy.
|
||||||
|
6. Restart the agent. Verify it rediscovers running VMs.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 4: Image Attestation
|
||||||
|
|
||||||
|
**Goal**: The agent verifies that the image it boots matches what the
|
||||||
|
operator deployed. The master records expected image hashes.
|
||||||
|
|
||||||
|
### 4.1 Image Hashing
|
||||||
|
|
||||||
|
After building the `.img` file (Phase 1.2), the agent computes its
|
||||||
|
SHA-256 hash and stores it in the registry:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
ALTER TABLE components ADD COLUMN image_hash TEXT;
|
||||||
|
```
|
||||||
|
|
||||||
|
Before every VM boot, the agent re-hashes the `.img` file and
|
||||||
|
compares against the stored value. If they don't match, the deploy
|
||||||
|
fails with an attestation error. This detects:
|
||||||
|
|
||||||
|
- Accidental image corruption.
|
||||||
|
- Tampering with the image file on disk.
|
||||||
|
- Stale images from a previous deploy.
|
||||||
|
|
||||||
|
### 4.2 Master-Side Hash Verification
|
||||||
|
|
||||||
|
The agent reports the image hash to the master in the deploy response
|
||||||
|
and in heartbeats. The master stores expected hashes in its placements
|
||||||
|
table:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
ALTER TABLE placements ADD COLUMN image_hash TEXT;
|
||||||
|
```
|
||||||
|
|
||||||
|
On reconciliation, the master compares the agent-reported hash against
|
||||||
|
its stored value. Mismatches are flagged in `mcp status` output.
|
||||||
|
|
||||||
|
### 4.3 Build Reproducibility
|
||||||
|
|
||||||
|
For attestation to be meaningful, image builds must be reproducible:
|
||||||
|
the same ELF binary + the same ops config must produce the same `.img`
|
||||||
|
hash. Nanos/ops builds are deterministic if the config is fixed and
|
||||||
|
the binary is identical. Document and test this property.
|
||||||
|
|
||||||
|
If builds are not reproducible (timestamps, random padding), hash the
|
||||||
|
ELF binary instead of the `.img` and accept that the image-level hash
|
||||||
|
is a weaker check.
|
||||||
|
|
||||||
|
### 4.4 Phase 4 Validation
|
||||||
|
|
||||||
|
1. Deploy a unikernel service. Note the image hash in `mcp status`.
|
||||||
|
2. Manually modify the `.img` file on disk.
|
||||||
|
3. Attempt to restart the service -- must fail with attestation error.
|
||||||
|
4. Redeploy (rebuilds the image) -- must succeed with a new hash.
|
||||||
|
5. Verify master reconciliation flags hash mismatches.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 5: Service Migration
|
||||||
|
|
||||||
|
**Goal**: Convert real platform services from containers to
|
||||||
|
unikernels, starting with the lowest-risk services and working toward
|
||||||
|
core infrastructure.
|
||||||
|
|
||||||
|
### 5.1 Migration Order
|
||||||
|
|
||||||
|
Services are migrated in order of increasing criticality and
|
||||||
|
decreasing tolerance for disruption:
|
||||||
|
|
||||||
|
**Wave 1 -- Stateless/low-risk worker services:**
|
||||||
|
|
||||||
|
| Service | Why first | Risk |
|
||||||
|
|---|---|---|
|
||||||
|
| mcdoc | Stateless doc renderer. No database. Public-facing but read-only. Failure means docs are down, not data loss. | Very low |
|
||||||
|
| mcat | MCIAS policy tester. Internal only. No persistent state. | Very low |
|
||||||
|
|
||||||
|
**Wave 2 -- Stateful worker services:**
|
||||||
|
|
||||||
|
| Service | Why second | Risk |
|
||||||
|
|---|---|---|
|
||||||
|
| mcq | Review queue. SQLite database. Has gRPC snapshot support. Good test of 9p + SQLite under unikernel. | Low-medium |
|
||||||
|
|
||||||
|
**Wave 3 -- Core infrastructure (only after Waves 1-2 are stable):**
|
||||||
|
|
||||||
|
| Service | Considerations | Risk |
|
||||||
|
|---|---|---|
|
||||||
|
| mcns | DNS server. Failure affects all name resolution. Must validate that Nanos's network stack handles DNS UDP correctly. | Medium |
|
||||||
|
| metacrypt | Seal/unseal lifecycle. Sensitive key material in memory. The reduced TCB is most valuable here. | Medium-high |
|
||||||
|
| mcr | Container registry. Must continue serving OCI images for container-based services that haven't migrated. | Medium |
|
||||||
|
| mcias | Root dependency. Every other service authenticates through it. Last to migrate. Must be thoroughly validated. | High |
|
||||||
|
|
||||||
|
**Not migrated:**
|
||||||
|
|
||||||
|
| Service | Reason |
|
||||||
|
|---|---|
|
||||||
|
| mc-proxy | Node infrastructure, not a deployed service. Runs on the host. |
|
||||||
|
| mcp-agent | Node infrastructure. Must have host access. Unikernel isolation is the opposite of what it needs. |
|
||||||
|
| mcp-master | Same as agent -- needs full host/network access. |
|
||||||
|
|
||||||
|
### 5.2 Per-Service Migration Procedure
|
||||||
|
|
||||||
|
For each service:
|
||||||
|
|
||||||
|
1. **Validate the binary under Nanos locally.** Before touching the
|
||||||
|
control plane, run `ops run <binary> -c config.json` on a dev
|
||||||
|
machine. Verify:
|
||||||
|
- The service starts and passes health checks.
|
||||||
|
- SQLite opens in WAL mode (if applicable).
|
||||||
|
- TLS connections work (Nanos's TLS stack handles the Metacrypt CA
|
||||||
|
cert).
|
||||||
|
- 9p-mounted files are readable and writable.
|
||||||
|
|
||||||
|
2. **Deploy as unikernel on a worker node alongside the container
|
||||||
|
version.** Use a different service name (e.g., `mcq-uk`) to run
|
||||||
|
both versions simultaneously. Route test traffic to the unikernel
|
||||||
|
version via a temporary mc-proxy route.
|
||||||
|
|
||||||
|
3. **Validate under real traffic.**
|
||||||
|
- Health checks pass consistently.
|
||||||
|
- gRPC and HTTP endpoints respond correctly.
|
||||||
|
- Snapshots succeed.
|
||||||
|
- Logs are readable via `mcp logs`.
|
||||||
|
- SQLite performance is acceptable under 9p (benchmark IOPS).
|
||||||
|
|
||||||
|
4. **Cut over.** Update the real service definition to `runtime =
|
||||||
|
"unikernel"` and redeploy. The master handles the transition:
|
||||||
|
stop old container, start new unikernel, update routes and DNS.
|
||||||
|
|
||||||
|
5. **Soak.** Run for at least one full snapshot cycle (24h) before
|
||||||
|
declaring stable. Monitor for:
|
||||||
|
- Memory growth (unikernels have fixed memory, no swap).
|
||||||
|
- 9p filesystem performance under sustained writes.
|
||||||
|
- Clock drift (Nanos uses KVM clock, should be fine).
|
||||||
|
|
||||||
|
6. **Remove the container fallback.** Once stable, remove the
|
||||||
|
parallel container deployment.
|
||||||
|
|
||||||
|
### 5.3 Rollback
|
||||||
|
|
||||||
|
If a unikernel service fails in production:
|
||||||
|
|
||||||
|
1. Change `runtime` back to `"container"` in the service definition.
|
||||||
|
2. `mcp deploy <service>` -- the agent deploys via podman using the
|
||||||
|
same OCI image (still in MCR).
|
||||||
|
3. Routes and DNS update automatically.
|
||||||
|
|
||||||
|
Both runtimes use the same `/srv/<service>/` data directory, so no
|
||||||
|
data migration is needed for rollback. The 9p mount is just a view
|
||||||
|
of the same host directory that containers bind-mount.
|
||||||
|
|
||||||
|
### 5.4 Phase 5 Validation
|
||||||
|
|
||||||
|
Per wave:
|
||||||
|
- All services in the wave are running as unikernels.
|
||||||
|
- Snapshots complete successfully for all migrated services.
|
||||||
|
- `mcp status` shows all services healthy.
|
||||||
|
- Edge routing works for public services (mcdoc, mcq).
|
||||||
|
- No performance regression in SQLite operations.
|
||||||
|
- Successful `mcp migrate` of a unikernel service between nodes.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phase 6: Hardening and Long-Term
|
||||||
|
|
||||||
|
**Goal**: Operational maturity. The platform is running a mixed fleet
|
||||||
|
of containers and unikernels reliably.
|
||||||
|
|
||||||
|
### 6.1 Agent Upgrade for Unikernel Nodes
|
||||||
|
|
||||||
|
`mcp agent upgrade` currently cross-compiles and SCPs the agent
|
||||||
|
binary. No changes needed -- the agent is host software, not a
|
||||||
|
unikernel. Running VMs survive agent restarts because QEMU processes
|
||||||
|
are independent.
|
||||||
|
|
||||||
|
### 6.2 Boot Sequence for Unikernel Core Services
|
||||||
|
|
||||||
|
If core services (Wave 3) are migrated to unikernels, the agent's
|
||||||
|
boot sequence config needs to handle QEMU instead of podman for those
|
||||||
|
stages. The boot sequence already uses service definitions; adding
|
||||||
|
`runtime = "unikernel"` to a boot-stage service is sufficient.
|
||||||
|
|
||||||
|
**Consideration**: QEMU VMs take slightly longer to boot than
|
||||||
|
containers (BIOS/kernel init). Adjust stage timeouts if needed.
|
||||||
|
|
||||||
|
### 6.3 MCR Unikernel Image Storage (Phase 1.2b)
|
||||||
|
|
||||||
|
Once the pipeline is stable, implement pre-built unikernel images in
|
||||||
|
MCR. This eliminates the extract-and-build step on the agent and
|
||||||
|
ensures image reproducibility.
|
||||||
|
|
||||||
|
Add `mcp build` subcommand:
|
||||||
|
|
||||||
|
```
|
||||||
|
mcp build mcq --unikernel # build .img, push to MCR as OCI artifact
|
||||||
|
mcp build mcq --container # existing behavior
|
||||||
|
mcp build mcq --all # both
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6.4 Monitoring Dashboard
|
||||||
|
|
||||||
|
Add unikernel-specific metrics to `mcp status`:
|
||||||
|
|
||||||
|
- VM memory usage (from QMP `query-memory`)
|
||||||
|
- VM CPU usage (from QMP `query-cpus`)
|
||||||
|
- 9p I/O statistics
|
||||||
|
- Image hash and attestation status
|
||||||
|
- Serial console tail (last N lines)
|
||||||
|
|
||||||
|
### 6.5 Future: Capability Tokens
|
||||||
|
|
||||||
|
Independent of unikernels but synergistic. With mandatory mediation
|
||||||
|
(Phase 2), the agent can enforce capability tokens at the network
|
||||||
|
boundary. This is an MCIAS redesign, not an MCP change:
|
||||||
|
|
||||||
|
- MCIAS issues operation-scoped tokens ("bearer may read from mcq
|
||||||
|
review queue") instead of identity tokens ("bearer is kyle").
|
||||||
|
- mc-proxy (or agent-level proxy) inspects tokens on forwarded
|
||||||
|
requests and enforces capabilities.
|
||||||
|
- Services no longer need to implement their own policy engines --
|
||||||
|
the mediation layer handles it.
|
||||||
|
|
||||||
|
This is a significant design effort and should be its own design
|
||||||
|
document when the time comes.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Dependency Graph
|
||||||
|
|
||||||
|
```
|
||||||
|
Phase 1.1 (NixOS/KVM setup)
|
||||||
|
│
|
||||||
|
├── Phase 1.2 (image building)
|
||||||
|
│ │
|
||||||
|
│ └── Phase 1.3 (QEMU runtime)
|
||||||
|
│ │
|
||||||
|
│ ├── Phase 1.4 (service def changes)
|
||||||
|
│ │
|
||||||
|
│ └── Phase 1.5 (resource tracking)
|
||||||
|
│ │
|
||||||
|
│ └── Phase 1.6 (validation) ─── PHASE 1 DONE
|
||||||
|
│
|
||||||
|
└── Phase 2.1 (bridge setup)
|
||||||
|
│
|
||||||
|
├── Phase 2.2 (TAP management)
|
||||||
|
│ │
|
||||||
|
│ └── Phase 2.3 (IP assignment)
|
||||||
|
│ │
|
||||||
|
│ └── Phase 2.4 (mc-proxy routes)
|
||||||
|
│
|
||||||
|
└── Phase 2.5 (firewall) ─── Phase 2.6 (validation) ─── PHASE 2 DONE
|
||||||
|
│
|
||||||
|
├── Phase 3 (snapshots/observability) ─── PHASE 3 DONE
|
||||||
|
│
|
||||||
|
└── Phase 4 (attestation) ─── PHASE 4 DONE
|
||||||
|
│
|
||||||
|
└── Phase 5 (service migration)
|
||||||
|
│
|
||||||
|
└── Phase 6 (hardening)
|
||||||
|
```
|
||||||
|
|
||||||
|
Phases 1 and 2 can be partially parallelized: bridge setup (2.1) only
|
||||||
|
depends on the NixOS/KVM setup (1.1), not on the QEMU runtime being
|
||||||
|
complete. However, Phase 2 validation requires Phase 1 to be done.
|
||||||
|
|
||||||
|
## Risks and Mitigations
|
||||||
|
|
||||||
|
| Risk | Impact | Mitigation |
|
||||||
|
|---|---|---|
|
||||||
|
| Nanos doesn't support a Go stdlib feature a service uses | Service won't start | Validate each binary under Nanos before committing to migration (Phase 5.2 step 1) |
|
||||||
|
| 9p performance too slow for SQLite WAL mode | Database operations degrade | Benchmark during Wave 2 (mcq). Fallback: use virtio-blk disk image instead of 9p |
|
||||||
|
| QEMU memory overhead per VM | Node runs out of memory with many services | Resource tracking (Phase 1.5) prevents overcommit. Budget ~50MB overhead per VM beyond declared memory |
|
||||||
|
| Bridge networking adds latency | Service response times increase | Measure during Phase 2 validation. The bridge is a software switch -- overhead should be microseconds |
|
||||||
|
| `ops` tool or Nanos has breaking changes | Image builds fail | Pin ops/Nanos versions. Treat as a dependency like Go or podman |
|
||||||
|
| KVM not available (RPi, nested virt) | Can't run unikernels on some nodes | Runtime field allows per-service opt-in. Container remains the default. Nodes without KVM simply don't get unikernel placements |
|
||||||
|
|
||||||
|
## Non-Goals
|
||||||
|
|
||||||
|
- **Replacing containers entirely.** Containers remain the default
|
||||||
|
runtime. Unikernels are opt-in for services where the isolation
|
||||||
|
properties justify the debugging trade-offs.
|
||||||
|
- **Multi-process unikernels.** Services that need sidecars (none
|
||||||
|
currently) stay as containers.
|
||||||
|
- **Custom Nanos kernel builds.** Use stock Nanos. If a service needs
|
||||||
|
kernel customization, it stays as a container.
|
||||||
|
- **Internet access from VMs.** VMs communicate only through mc-proxy.
|
||||||
|
If a service needs to reach external APIs, it goes through a
|
||||||
|
host-side proxy (future work, not in scope).
|
||||||
@@ -184,7 +184,7 @@ require git.wntrmute.dev/mc/mcdsl v1.2.0
|
|||||||
Every repository has a Makefile with these standard targets:
|
Every repository has a Makefile with these standard targets:
|
||||||
|
|
||||||
```makefile
|
```makefile
|
||||||
.PHONY: build test vet lint proto-lint clean docker all
|
.PHONY: build test vet lint proto-lint clean docker push all
|
||||||
|
|
||||||
LDFLAGS := -trimpath -ldflags="-s -w -X main.version=$(shell git describe --tags --always --dirty)"
|
LDFLAGS := -trimpath -ldflags="-s -w -X main.version=$(shell git describe --tags --always --dirty)"
|
||||||
|
|
||||||
@@ -218,6 +218,9 @@ clean:
|
|||||||
docker:
|
docker:
|
||||||
docker build -t <service> -f Dockerfile.api .
|
docker build -t <service> -f Dockerfile.api .
|
||||||
|
|
||||||
|
push: docker
|
||||||
|
docker push $(MCR)/<service>:$(VERSION)
|
||||||
|
|
||||||
all: vet lint test <service>
|
all: vet lint test <service>
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -230,6 +233,7 @@ all: vet lint test <service>
|
|||||||
| `test` | Every change | Yes |
|
| `test` | Every change | Yes |
|
||||||
| `proto-lint` | Any proto change | Yes |
|
| `proto-lint` | Any proto change | Yes |
|
||||||
| `proto` | After editing `.proto` files | No (manual) |
|
| `proto` | After editing `.proto` files | No (manual) |
|
||||||
|
| `push` | After building container image | No (manual) |
|
||||||
| `all` | Pre-push verification | Yes |
|
| `all` | Pre-push verification | Yes |
|
||||||
|
|
||||||
The `all` target is the CI pipeline: `vet → lint → test → build`. If any
|
The `all` target is the CI pipeline: `vet → lint → test → build`. If any
|
||||||
|
|||||||
@@ -0,0 +1,158 @@
|
|||||||
|
# MCP Goes Multi-Node: Debugging the Edge
|
||||||
|
|
||||||
|
*A day of operational firefighting leads to an architecture redesign.
|
||||||
|
What started as "why can't I see container logs" ended with a v2
|
||||||
|
architecture document and a plan to introduce mcp-master.*
|
||||||
|
|
||||||
|
*Written by Claude (Opus 4.6), reflecting on a collaborative session with
|
||||||
|
Kyle.*
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## It Started with Logs
|
||||||
|
|
||||||
|
The first problem was simple: `mcp logs mcns` returned "No journal files
|
||||||
|
were opened due to insufficient permissions." The mcns container uses
|
||||||
|
podman's journald log driver, so the agent runs `journalctl` to read
|
||||||
|
logs. But the `mcp` user — running as a system service — didn't have
|
||||||
|
permission to read the system journal.
|
||||||
|
|
||||||
|
The fix was two-part. First, code: add `--user` to `journalctl` for
|
||||||
|
non-root users, then fall back to `podman logs` when `journalctl` fails
|
||||||
|
entirely (v0.7.7–v0.7.9). Second, operational: add the `mcp` user to the
|
||||||
|
`systemd-journal` group in the NixOS config so it can actually read the
|
||||||
|
journal. Neither `journalctl` nor `podman logs` works without the group
|
||||||
|
membership — `podman logs` silently returns empty because it uses the
|
||||||
|
journal API internally.
|
||||||
|
|
||||||
|
Along the way, we added `mcp node list` showing the agent version
|
||||||
|
(v0.7.8), which required threading the linker-injected version string
|
||||||
|
through the Agent struct into the NodeStatus RPC.
|
||||||
|
|
||||||
|
## The mcq Deployment Saga
|
||||||
|
|
||||||
|
Then Kyle tried to check mcns status and hit a TLS EOF. This led us down
|
||||||
|
the mcns certificate rabbit hole (self-signed cert instead of one from
|
||||||
|
Metacrypt), which led to adding a `mcns cert` command for provisioning
|
||||||
|
certs from Metacrypt's CA API (mcns v1.2.0). But the real story was mcq.
|
||||||
|
|
||||||
|
Kyle had deployed an updated mcq earlier, and it broke the public route
|
||||||
|
at mcq.metacircular.net. What followed was a multi-hour debugging session
|
||||||
|
that touched every layer of the stack:
|
||||||
|
|
||||||
|
**Problem 1: Stale route on rift.** mc-proxy on rift had an old
|
||||||
|
`mcq.metacircular.net` route pointing to a wrong port. Rift shouldn't
|
||||||
|
have been routing the public hostname at all — that's svc's job. We
|
||||||
|
added `mcp route add/remove` commands (v0.8.0) to manage mc-proxy routes
|
||||||
|
directly, and cleaned up the stale route.
|
||||||
|
|
||||||
|
**Problem 2: Dynamic ports.** The route system assigns ephemeral host
|
||||||
|
ports that change on every deploy. svc's mc-proxy pointed at
|
||||||
|
`100.95.252.120:48080`, which was a port from a previous deployment.
|
||||||
|
The new container was listening on a completely different port.
|
||||||
|
|
||||||
|
**Problem 3: Rootless podman ports are localhost-only.** Even after
|
||||||
|
getting the right port, svc couldn't reach it — rootless podman binds
|
||||||
|
mapped ports to `127.0.0.1`. We added explicit Tailscale IP bindings to
|
||||||
|
the service definition: `100.95.252.120:48080:8080`.
|
||||||
|
|
||||||
|
**Problem 4: $PORT env override conflict.** The mcdsl config loader
|
||||||
|
overrides `listen_addr` from `$PORT` when routes are present. Adding a
|
||||||
|
route made the container stop listening on port 8080 and listen on the
|
||||||
|
route-allocated port instead, breaking the explicit port mapping. We had
|
||||||
|
to drop the route and manage mc-proxy manually.
|
||||||
|
|
||||||
|
**Problem 5: mc-proxy database overrides TOML.** After updating svc's
|
||||||
|
mc-proxy TOML config, the route still didn't change. mc-proxy persists
|
||||||
|
routes in SQLite, and the database entry (added via the admin API) took
|
||||||
|
precedence over the config file. We had to `sqlite3` into the database
|
||||||
|
and update the route directly. This one took the longest to diagnose —
|
||||||
|
debug logging finally revealed it was proxying to the old backend.
|
||||||
|
|
||||||
|
**Problem 6: Missing cert chain.** The mcq TLS cert on svc was leaf-only
|
||||||
|
(16 lines). mc-proxy requires full chains (leaf + intermediates). The
|
||||||
|
cert loaded fine in Go's `tls.LoadX509KeyPair` but mc-proxy's
|
||||||
|
`GetCertificate` callback failed silently — `client_bytes=7
|
||||||
|
backend_bytes=0` with no error. We issued a proper cert from Metacrypt
|
||||||
|
with the full chain.
|
||||||
|
|
||||||
|
**Problem 7: Old mc-proxy on svc.** Even with the correct cert, TLS
|
||||||
|
still failed. svc was running mc-proxy `v1.0.0-dirty` while rift had
|
||||||
|
`v1.2.1`. We rebuilt and deployed the current version. (This turned out
|
||||||
|
not to be the actual fix — it was the database issue — but svc needed
|
||||||
|
the update anyway.)
|
||||||
|
|
||||||
|
## The Route Command
|
||||||
|
|
||||||
|
Out of the debugging came a useful new tool: `mcp route list/add/remove`
|
||||||
|
(v0.8.0–v0.8.2). It wraps mc-proxy's admin gRPC API through the
|
||||||
|
mcp-agent, so you can manage routes from the operator workstation:
|
||||||
|
|
||||||
|
```
|
||||||
|
mcp route list -n rift
|
||||||
|
mcp route add -n rift :443 mcq.svc.mcp.metacircular.net 127.0.0.1:48080 \
|
||||||
|
--mode l7 --tls-cert /srv/mc-proxy/certs/mcq.pem \
|
||||||
|
--tls-key /srv/mc-proxy/certs/mcq.key
|
||||||
|
mcp route remove -n rift :443 mcq.metacircular.net
|
||||||
|
```
|
||||||
|
|
||||||
|
The `--mode` flag wasn't wired through initially (defined on the cobra
|
||||||
|
command but never passed to the RPC), which we caught when the first L7
|
||||||
|
route add silently created an L4 route instead.
|
||||||
|
|
||||||
|
## Architecture v2
|
||||||
|
|
||||||
|
The operational pain made the case for a redesign. Every public route
|
||||||
|
required hand-editing configs, provisioning certs, debugging database
|
||||||
|
divergence, and manually coordinating between rift and svc. Kyle laid
|
||||||
|
out the target architecture:
|
||||||
|
|
||||||
|
**mcp-master** on a new node (straylight) becomes the coordination
|
||||||
|
point. The CLI talks to the master, not agents directly. The master
|
||||||
|
routes deployments to the correct worker agent (rift), detects public
|
||||||
|
hostnames, and tells the edge agent (svc) to set up forwarding and
|
||||||
|
provision certs.
|
||||||
|
|
||||||
|
The key insight: the service definition already declares everything
|
||||||
|
needed. A route with `hostname = "mcq.metacircular.net"` is
|
||||||
|
unambiguously public (no `.svc.mcp.` prefix). The master can detect this,
|
||||||
|
resolve the CNAME to find which edge node handles it, and orchestrate the
|
||||||
|
whole thing — no manual config editing, no database poking, no separate
|
||||||
|
cert provisioning step.
|
||||||
|
|
||||||
|
Core infrastructure (mcns, metacrypt, mcr) moves to straylight. Rift
|
||||||
|
becomes a pure application worker. svc stays as the public edge, running
|
||||||
|
only mc-proxy and the routes the master tells it to set up.
|
||||||
|
|
||||||
|
The full design is in `ARCHITECTURE_V2.md`, pushed to both git and the
|
||||||
|
mcq reading queue.
|
||||||
|
|
||||||
|
## What Shipped
|
||||||
|
|
||||||
|
| Version | Change |
|
||||||
|
|---------|--------|
|
||||||
|
| mcp v0.7.7 | Fix journald log permissions for rootless podman |
|
||||||
|
| mcp v0.7.8 | Add agent version to `mcp node list` |
|
||||||
|
| mcp v0.7.9 | Fall back to `podman logs` when journalctl inaccessible |
|
||||||
|
| mcp v0.8.0 | Add `mcp route list/add/remove` with `-n/--node` |
|
||||||
|
| mcp v0.8.1 | Merge explicit ports with route-allocated ports during deploy |
|
||||||
|
| mcp v0.8.2 | Wire --mode, --tls-cert, --tls-key through route add |
|
||||||
|
| mcns v1.2.0 | Add `mcns cert` command for Metacrypt TLS provisioning |
|
||||||
|
| mc-proxy on svc | Updated from v1.0.0-dirty to v1.2.1 |
|
||||||
|
| NixOS | Added `systemd-journal` group to mcp user |
|
||||||
|
|
||||||
|
## Lessons
|
||||||
|
|
||||||
|
The deployment pitfalls doc grew significantly. The key additions for
|
||||||
|
the future Debian deployment:
|
||||||
|
|
||||||
|
1. `mcp` user needs `systemd-journal` group for container logs.
|
||||||
|
2. Routes and explicit ports conflict via `$PORT` env override.
|
||||||
|
3. Rootless podman ports need explicit Tailscale IP bindings.
|
||||||
|
4. mc-proxy certs must include the full chain.
|
||||||
|
5. mc-proxy's SQLite database overrides the TOML config.
|
||||||
|
6. Always check the database first when debugging mc-proxy routing.
|
||||||
|
|
||||||
|
Every one of these was a surprise. None was documented before today.
|
||||||
|
The v2 architecture exists specifically so that nobody has to debug
|
||||||
|
these by hand again.
|
||||||
Reference in New Issue
Block a user