Container isolation: fix stdio stall + per-exec connection leak; add guest image pipeline

Host-side (ships with a normal swift build):
- LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking)
  so a wedged stream can't head-of-line-block sibling execs' relays; atomic
  stdio-or-abort start (patches #5, #6).
- Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged
  channel (patch #7).
- ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on
  force-close — fixes a per-turn leak (per-exec vsock/gRPC connection +
  runConnections() task) in the long-lived shared control container. Likely
  the "degrades until app restart" root cause.
- ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so
  a failed launch settles as retryable instead of locking the composer.

Guest-side (rides the custom vminitd initfs; inert until the image is built):
- ManagedProcess: offload the blocking start off the gRPC event loop (patch #8).
- Per-exec cgroups (patch #9) recorded as design only — cross-cutting.

Pipeline:
- .github/workflows/vminit-image.yml builds vminitd from the vendored source
  and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed
  at the custom image.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
2026-07-13 18:46:48 -07:00
co-authored by Claude Opus 4.8
parent 4bcf34d91a
commit 7972dfb02d
4 changed files with 188 additions and 23 deletions
+74 -2
View File
@@ -41,6 +41,72 @@ in-tree means the patch can't be lost to a dependency re-resolve.
were dropped, and the corresponding `.testTarget(...)` entries removed from `Package.swift`. The
library/executable targets we build are untouched.
5. **`Sources/Containerization/LinuxProcess.swift` — non-blocking stdio relay.**
Upstream's `setupIO` relays guest stdout/stderr with `FileHandle.availableData`, a **blocking**
read, from inside a `readabilityHandler`. Those handlers run on Foundation's shared readability
queue, so if one exec's guest stdout wedged mid-stream that blocking read parked the shared thread
and head-of-line-blocked **every** other exec's stdout/stderr relay across all containers — one
stuck session froze the others. The patch marks each connected fd `O_NONBLOCK` and drains it via a
new `nucleicDrainNonBlocking` (returns bytes + EOF, never blocks; EAGAIN just waits for the next
readable event). A wedged stream is now contained to its own exec. Marked `[Nucleic vendored patch]`
(the two static helpers `nucleicSetNonBlocking`/`nucleicDrainNonBlocking` and the two rewritten
`readabilityHandler` blocks). Requires host-side POSIX `read`/`fcntl`/`errno`.
6. **`Sources/Containerization/LinuxProcess.swift` — atomic stdio-or-abort start.**
In `start()`, after `setupIO` returns, if a *configured* stdio stream never connected from the
guest (its `FileHandle` is nil — patch #3's logged failure), the patch tears the just-created exec
back down (`agent.deleteProcess`) and throws instead of calling `startProcess`. Upstream proceeds
and runs a process with a dead stream (stdin never delivered → hangs; stdout never read → the "no
output, just a spinner" 60s stall in Nucleic Control). Now that permanent silent stall surfaces as
a clean, retryable start error. Marked `[Nucleic vendored patch]` (the guard block before
`startProcess`).
7. **`Sources/Containerization/Vminitd.swift` — bounded teardown RPC.**
`deleteProcess` now sends a 30s `CallOptions.timeout` (upstream sends none, so it can block
forever on a wedged agent channel). Nucleic calls `LinuxProcess.delete()` after every turn to
reclaim the per-exec vsock/gRPC connection `exec()` dials; an unbounded `deleteProcess` would let
that reclaim hang and the connection leak. On the thrown deadline, `performDeletion` still closes
the agent connection. Marked `[Nucleic vendored patch]` (the `callOpts` block in `deleteProcess`).
NOTE: this pairs with a Nucleic-side change in `ContainerizedProcessHandle` (call `delete()` after
the exec exits / on force-close) — without that caller, upstream never deletes execs at all and
the shared control container leaks a connection + `runConnections()` task per turn.
### GUEST-side patches (require rebuilding the initfs — see below)
Patches #1–#7 are host-side (the `Containerization` library), shipped by a normal `swift build`.
Patches #8+ live in `vminitd/` (the guest agent), which rides in the initfs OCI image. They are INERT
until that image is rebuilt from this source and published, and `ContainerEngine.vminitReference`
points at it. That is now automated: **`.github/workflows/vminit-image.yml`** builds vminitd from this
vendored tree and pushes `ghcr.io/abkslm/vminit:<tag>`; `vminitReference` is pinned to that custom
image. Bump the `-nucleicN` tag suffix and re-run the workflow whenever a guest patch changes.
8. **`vminitd/Sources/VminitdCore/ManagedProcess.swift` — offload the blocking start off the event loop.**
`ManagedProcess.start()` did synchronous, potentially slow pipe reads (waiting for `vmexec` to
return the pid, then for the error pipe to close) while holding `state`'s Mutex, ON the calling
task — which is the gRPC handler's event-loop thread. A slow start therefore parked the loop and
head-of-line-blocked sibling execs' control RPCs sharing it. The patch splits the body into a
synchronous `startBlocking()` and an async `start()` that runs it on `DispatchQueue.global` via a
checked continuation, keeping the loop responsive. Safe because the body has no `await` and
`ManagedProcess` is `Sendable`. Marked `[Nucleic vendored patch]`.
### PLANNED guest patch (design recorded; NOT yet implemented)
9. **Per-exec cgroups (memory/cpu/pids isolation).** Today the whole container shares ONE cgroup
(`/container/<id>`): `vmexec run` places the init there via the OCI `cgroupsPath` + `applyResources`
(`RunCommand.swift`), and each exec joins it via `loadFromPid(init.pid).addProcess` in
`ManagedProcess.start`. So one session's runaway RSS trips the VM OOM-killer against a *random*
sibling. Target layout (cgroup v2): make `/container/<id>` an intermediary (enable
`cgroup.subtree_control` — `Cgroup2Manager.toggleSubtreeControllers` already skips the leaf so this
composes), move init to a leaf `/container/<id>/init`, and place each exec in its own leaf
`/container/<id>/<execID>` with generous `memory.high`/`memory.max`/`cpu.max`/`pids.max` so a
runaway session is throttled/OOM-killed *within its own cgroup*, siblings untouched — WITHOUT
hard-partitioning RAM (soft limits preserve burst). This is CROSS-CUTTING, not a one-file patch:
the per-exec limits must be carried on the exec RPC (the `CreateProcess`/exec OCI spec has no
resources field today), which means a protobuf field (`SandboxContext`) + host-side plumbing
(`Vminitd.createProcess` / `ContainerEngine.exec`) in addition to the vminitd cgroup restructure
(`ManagedContainer`, `ManagedProcess`, `vmexec/RunCommand`). Sequence it after #8 lands via CI, and
validate in a real container (a wrong v2 hierarchy fails at runtime, not at compile).
## Re-vendoring a newer upstream commit
1. `git clone` upstream (or copy `.build/checkouts/containerization` after bumping the URL pin
@@ -49,7 +115,13 @@ in-tree means the patch can't be lost to a dependency re-resolve.
--exclude=images/ <upstream>/ third_party/containerization/`
3. Remove the `.testTarget(...)` blocks from `third_party/containerization/Package.swift`.
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward), patch #2
(`LinuxProcess.killProcessGroup(_:)`), and patch #3 (the `setupIO` stdio-connection log + its
`import os` / `nucleicIOLog`). Grep for `[Nucleic vendored patch]` to find every site.
(`LinuxProcess.killProcessGroup(_:)`), patch #3 (the `setupIO` stdio-connection log + its
`import os` / `nucleicIOLog`), patch #5 (the non-blocking stdio relay: `nucleicSetNonBlocking` /
`nucleicDrainNonBlocking` + the rewritten `readabilityHandler` blocks), and patch #6 (the atomic
stdio-or-abort guard in `start()`), patch #7 (the bounded `deleteProcess` timeout in
`Vminitd.swift`), and patch #8 (the `ManagedProcess.start` event-loop offload in `vminitd/`). Grep
for `[Nucleic vendored patch]` to find every site. Patch #9 (per-exec cgroups) is design-only so
far — see its entry. After re-applying any `vminitd/` patch, re-run `.github/workflows/vminit-image.yml`
to rebuild + publish the custom init image, and bump `ContainerEngine.vminitReference`.
5. Update the commit hash above and in the root `Package.swift` comment.
6. `swift build` and run the balloon tests.