Files
containerization/PATCHES.md
T

151 lines
12 KiB
Markdown
Raw Normal View History

# Vendored `containerization` — Nucleic patches
This is a **vendored copy** of [apple/containerization](https://github.com/apple/containerization)
at upstream commit `6b7b42ca3efeee8c706070e4355e6a807c5336ae`, referenced by the root `Package.swift`
via `.package(path: "third_party/containerization")` instead of the github URL.
It is vendored (not pulled) because we carry a local patch upstream doesn't have. Keeping it
in-tree means the patch can't be lost to a dependency re-resolve.
## What's changed vs. upstream
1. **`Sources/Containerization/LinuxContainer.swift` — forward VM extensions.**
`LinuxContainer.Configuration` gains a `vmExtensions: [any Sendable]` field, and
`LinuxContainer` assigns it into `VMConfiguration.extensions` when it builds the VM config.
Upstream already supports `VMConfiguration.extensions` + the `VZInstanceExtension` hook
(`configureVZ`/`didCreate`), but `LinuxContainer` — the only entry point we use — never forwarded
it, so there was no way to attach a device (e.g. a virtio memory balloon) to a container's VM.
Search for the marker comment `[Nucleic vendored patch]` to find both edit sites.
Nucleic uses this to attach a `VZVirtioTraditionalMemoryBalloonDeviceConfiguration` and drive its
target at runtime for automatic VM memory reclamation — see `MemoryBalloon.swift` /
`ContainerEngine` in NucleicCore.
2. **`Sources/Containerization/LinuxProcess.swift` — process-group kill.**
`LinuxProcess` gains `killProcessGroup(_:)`, which signals the negative pid (`-pid`) so the
guest's `kill(2)` targets the exec'd process's whole **process group**, not just the leader.
Every exec is `setsid()`'d by `vmexec`, so the process is its own group leader (pgid == pid) and
a group signal reaches the children it forked. Upstream only exposes the leader-only `kill(_:)`,
which let a forked child survive a Stop in a long-lived shared container. Marked with
`[Nucleic vendored patch]`; used by `ContainerizedProcessHandle.sendSignal` in NucleicCore.
3. **`Sources/Containerization/LinuxProcess.swift` — stdio-connection diagnostics (log-only).**
`setupIO` logs (`os.Logger`, subsystem `com.nucleic`, category `container-io`) when a *configured*
stdio stream's guest side never connects — which leaves its host `FileHandle` nil, so the relay /
readability handler is never wired and the agent's stdin is never delivered (it hangs) or its
stdout is never read (the "no output, just a spinner" symptom in Nucleic Control containers).
Behavior is unchanged; it only surfaces the failing stream. Marked `[Nucleic vendored patch]`
(the `import os`, the `nucleicIOLog` static, and the per-stream check in `setupIO`). All three are
wrapped in `#if canImport(os)` — non-Xcode toolchains (e.g. a swiftly Swift used to cross-build the
host framework) resolve Foundation/Virtualization but not the `os` overlay, so the diagnostic
degrades to a no-op there instead of failing the build; Xcode (the local `make vminit-image` path)
builds keep it.
4. **Trimmed for footprint (no behavior change).** `Tests/`, `docs/`, `examples/`, and `images/`
were dropped, and the corresponding `.testTarget(...)` entries removed from `Package.swift`. The
library/executable targets we build are untouched.
5. **`Sources/Containerization/LinuxProcess.swift` — non-blocking stdio relay.**
Upstream's `setupIO` relays guest stdout/stderr with `FileHandle.availableData`, a **blocking**
read, from inside a `readabilityHandler`. Those handlers run on Foundation's shared readability
queue, so if one exec's guest stdout wedged mid-stream that blocking read parked the shared thread
and head-of-line-blocked **every** other exec's stdout/stderr relay across all containers — one
stuck session froze the others. The patch marks each connected fd `O_NONBLOCK` and drains it via a
new `nucleicDrainNonBlocking` (returns bytes + EOF, never blocks; EAGAIN just waits for the next
readable event). A wedged stream is now contained to its own exec. Marked `[Nucleic vendored patch]`
(the two static helpers `nucleicSetNonBlocking`/`nucleicDrainNonBlocking` and the two rewritten
`readabilityHandler` blocks). Requires host-side POSIX `read`/`fcntl`/`errno`.
6. **`Sources/Containerization/LinuxProcess.swift` — atomic stdio-or-abort start.**
In `start()`, after `setupIO` returns, if a *configured* stdio stream never connected from the
guest (its `FileHandle` is nil — patch #3's logged failure), the patch tears the just-created exec
back down (`agent.deleteProcess`) and throws instead of calling `startProcess`. Upstream proceeds
and runs a process with a dead stream (stdin never delivered → hangs; stdout never read → the "no
output, just a spinner" 60s stall in Nucleic Control). Now that permanent silent stall surfaces as
a clean, retryable start error. Marked `[Nucleic vendored patch]` (the guard block before
`startProcess`).
7. **`Sources/Containerization/Vminitd.swift` — bounded teardown RPC.**
`deleteProcess` now sends a 30s `CallOptions.timeout` (upstream sends none, so it can block
forever on a wedged agent channel). Nucleic calls `LinuxProcess.delete()` after every turn to
reclaim the per-exec vsock/gRPC connection `exec()` dials; an unbounded `deleteProcess` would let
that reclaim hang and the connection leak. On the thrown deadline, `performDeletion` still closes
the agent connection. Marked `[Nucleic vendored patch]` (the `callOpts` block in `deleteProcess`).
NOTE: this pairs with a Nucleic-side change in `ContainerizedProcessHandle` (call `delete()` after
the exec exits / on force-close) — without that caller, upstream never deletes execs at all and
the shared control container leaks a connection + `runConnections()` task per turn.
### GUEST-side patches (require rebuilding the initfs — see below)
Patches #1–#7 are host-side (the `Containerization` library), shipped by a normal `swift build`.
Patches #8+ live in `vminitd/` (the guest agent), which rides in the initfs OCI image. They are INERT
until that image is rebuilt from this source and published, and `ContainerEngine.vminitReference`
points at it. Build it with **`make vminit-image`** (root Makefile) — it builds cctl + the guest
vminitd/vmexec from this vendored tree and packages `ghcr.io/abkslm/vminit:<tag>` into the local cctl
store; `make vminit-image-push` publishes it (authenticate once with `make vminit-image-login`, which
stores a GHCR token in the macOS Keychain — or set `REGISTRY_HOST`/`USERNAME`/`TOKEN`), and
`vminitReference` is pinned to that custom image. First time on a machine, run `make vminit-image-prep` once (installs
the swiftly toolchain + musl SDK the guest cross-build needs). Bump the `-nucleicN` tag suffix and
rebuild whenever a guest patch changes. Built locally, not in CI: the host framework needs the macOS
26+ Virtualization SDK that GitHub-hosted runners lack.
8. **`vminitd/Sources/VminitdCore/ManagedProcess.swift` — offload the blocking start off the event loop.**
`ManagedProcess.start()` did synchronous, potentially slow pipe reads (waiting for `vmexec` to
return the pid, then for the error pipe to close) while holding `state`'s Mutex, ON the calling
task — which is the gRPC handler's event-loop thread. A slow start therefore parked the loop and
head-of-line-blocked sibling execs' control RPCs sharing it. The patch splits the body into a
synchronous `startBlocking()` and an async `start()` that runs it on `DispatchQueue.global` via a
checked continuation, keeping the loop responsive. Safe because the body has no `await` and
`ManagedProcess` is `Sendable`. Marked `[Nucleic vendored patch]`.
9. **Per-exec cgroups (OOM/CPU/pids isolation).** Upstream puts the container init AND every exec in
ONE cgroup (`/container/<id>`), so one session's runaway RSS trips the in-VM OOM-killer against a
*random* sibling, and a fork bomb / CPU hog hits the whole box. This patch makes `/container/<id>`
an intermediary: the resource ceiling stays on it, `ManagedContainer.init` moves the init into its
own leaf (`/container/<id>/init`) which enables `cgroup.subtree_control` up the chain, and
`ManagedProcess.start` places each exec in its OWN child (`/container/<id>/<execID>`) with
`memory.oom.group=1` (a runaway session's OOM kills only *its* tree), a fair `cpu.weight`, and a
`pids.max` fork-bomb backstop. New `Cgroup2Manager` helpers: `setOomGroup`/`setCpuWeight`/
`setPidsMax`/`remove`. **Best-effort with a graceful fallback**: if any step of the per-exec setup
fails it wipes the partial state and reverts to the flat layout, and `ManagedProcess` falls back to
the container cgroup per exec — so a cgroup hiccup degrades to today's behavior, never a failed
start. `ManagedContainer.execCgroupParent == nil` marks flat mode.
Beyond scoped-OOM, a **hard host-configured per-exec `memory.max`** is also wired — WITHOUT a
protobuf change, because the exec already ships the full OCI `Spec` and the guest just ignored
`linux.resources`. Host: `LinuxProcessConfiguration.memoryLimitInBytes` → `LinuxContainer.exec`
stamps it onto `spec.linux.resources.memory.limit`. Guest: `Server+GRPC.createProcess` reads that
back and passes it to `createExec`/`ManagedProcess`, which sets `memory.max` (new
`Cgroup2Manager.setMemoryMax`) on the exec's cgroup — so a session can't consume the whole box
before its own OOM. Driven by Nucleic's `ContainerServiceSettings.controlPerSessionMemoryGiB`
(default 0 = off; applied only to the shared control container, via `ContainerManager.exec`), so the
default stays scoped-OOM-only. Marked `[Nucleic vendored patch]` across `Cgroup2Manager.swift`,
`ManagedContainer.swift`, `ManagedProcess.swift`, `Server+GRPC.swift` (guest) and
`LinuxProcessConfiguration.swift`, `LinuxContainer.swift` (host). **COMPILE-VERIFIED (host + musl
cross-build); NOT yet runtime-validated** — a
wrong cgroup-v2 hierarchy fails at runtime, so boot a container with the new image and confirm
sessions start, `/sys/fs/cgroup/container/<id>/<execID>` exists per session, and a hog is contained,
before pointing a shipping build at it. `vmexec/RunCommand` is unchanged: it still applies
`linux.resources` at `linux.cgroupsPath`, which the patch repoints (init leaf) and clears
accordingly.
## Re-vendoring a newer upstream commit
1. `git clone` upstream (or copy `.build/checkouts/containerization` after bumping the URL pin
temporarily), check out the desired commit.
2. `rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \
--exclude=images/ <upstream>/ third_party/containerization/`
3. Remove the `.testTarget(...)` blocks from `third_party/containerization/Package.swift`.
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward), patch #2
(`LinuxProcess.killProcessGroup(_:)`), patch #3 (the `setupIO` stdio-connection log + its
`import os` / `nucleicIOLog`), patch #5 (the non-blocking stdio relay: `nucleicSetNonBlocking` /
`nucleicDrainNonBlocking` + the rewritten `readabilityHandler` blocks), and patch #6 (the atomic
stdio-or-abort guard in `start()`), patch #7 (the bounded `deleteProcess` timeout in
`Vminitd.swift`), and patch #8 (the `ManagedProcess.start` event-loop offload in `vminitd/`). Grep
for `[Nucleic vendored patch]` to find every site, and patch #9 (per-exec cgroups) across
`Cgroup2Manager.swift` / `ManagedContainer.swift` / `ManagedProcess.swift`. After re-applying any
`vminitd/` patch, rebuild + publish the custom init image
with `make vminit-image` + `make vminit-image-push`, and bump `ContainerEngine.vminitReference`.
5. Update the commit hash above and in the root `Package.swift` comment.
6. `swift build` and run the balloon tests.