Containerized control sessions could hang with no output ("Working…" forever): the agent's exec
stdio over the vminitd vsock channel intermittently failed to carry bytes, so claude ran and
reached the approval server but its stdin/stdout never connected — it idled in interactive
stream-json mode and never exited. Forensics (nucleic.sqlite + per-session claude-home MCP logs)
showed every stdio layer byte-identical to a working state, i.e. a flaky framework race, not a
regression in our code. Make the failure recoverable and visible instead of an eternal spinner:
- ClaudeCodeBackend: a 60s spawn watchdog on containerized runs — no first stdout → emit a
recoverable error (with guest stderr + container probe), SIGKILL the wedged process, and finish
the run errored, instead of awaiting stdoutLines forever.
- Stop reliability: ProcessHandle.forceCloseStreams() (ContainerizedProcessHandle finishes its
line streams host-side; default no-op for the host pipe handle), wired into every kill
escalation (terminate/interruptThenKill/killGroupAfter) so a force-killed run always settles
even when the guest wait/stdio RPC wedges — the real cause of "Stop is inconsistent".
- Cap MCP_TIMEOUT (connection) to 60s so an unreachable approval server can't wedge startup for
~24.8 days; MCP_TOOL_TIMEOUT stays unbounded for human-answered approvals.
- LinuxProcess.setupIO logs which stdio stream fails to connect (os.Logger, com.nucleic /
container-io) so a stall pinpoints the failing stream. Vendored patch #3.
- Tests: stop-escalation force-close, responsive-process no-op, MCP timeout asymmetry.
Co-Authored-By: Claude Opus 4.8 <[email protected]>
3.7 KiB
Vendored containerization — Nucleic patches
This is a vendored copy of apple/containerization
at upstream commit 6b7b42ca3efeee8c706070e4355e6a807c5336ae, referenced by the root Package.swift
via .package(path: "third_party/containerization") instead of the github URL.
It is vendored (not pulled) because we carry a local patch upstream doesn't have. Keeping it in-tree means the patch can't be lost to a dependency re-resolve.
What's changed vs. upstream
-
Sources/Containerization/LinuxContainer.swift— forward VM extensions.LinuxContainer.Configurationgains avmExtensions: [any Sendable]field, andLinuxContainerassigns it intoVMConfiguration.extensionswhen it builds the VM config. Upstream already supportsVMConfiguration.extensions+ theVZInstanceExtensionhook (configureVZ/didCreate), butLinuxContainer— the only entry point we use — never forwarded it, so there was no way to attach a device (e.g. a virtio memory balloon) to a container's VM. Search for the marker comment[Nucleic vendored patch]to find both edit sites.Nucleic uses this to attach a
VZVirtioTraditionalMemoryBalloonDeviceConfigurationand drive its target at runtime for automatic VM memory reclamation — seeMemoryBalloon.swift/ContainerEnginein NucleicCore. -
Sources/Containerization/LinuxProcess.swift— process-group kill.LinuxProcessgainskillProcessGroup(_:), which signals the negative pid (-pid) so the guest'skill(2)targets the exec'd process's whole process group, not just the leader. Every exec issetsid()'d byvmexec, so the process is its own group leader (pgid == pid) and a group signal reaches the children it forked. Upstream only exposes the leader-onlykill(_:), which let a forked child survive a Stop in a long-lived shared container. Marked with[Nucleic vendored patch]; used byContainerizedProcessHandle.sendSignalin NucleicCore. -
Sources/Containerization/LinuxProcess.swift— stdio-connection diagnostics (log-only).setupIOlogs (os.Logger, subsystemcom.nucleic, categorycontainer-io) when a configured stdio stream's guest side never connects — which leaves its hostFileHandlenil, so the relay / readability handler is never wired and the agent's stdin is never delivered (it hangs) or its stdout is never read (the "no output, just a spinner" symptom in Nucleic Control containers). Behavior is unchanged; it only surfaces the failing stream. Marked[Nucleic vendored patch](theimport os, thenucleicIOLogstatic, and the per-stream check insetupIO). -
Trimmed for footprint (no behavior change).
Tests/,docs/,examples/, andimages/were dropped, and the corresponding.testTarget(...)entries removed fromPackage.swift. The library/executable targets we build are untouched.
Re-vendoring a newer upstream commit
git cloneupstream (or copy.build/checkouts/containerizationafter bumping the URL pin temporarily), check out the desired commit.rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \ --exclude=images/ <upstream>/ third_party/containerization/- Remove the
.testTarget(...)blocks fromthird_party/containerization/Package.swift. - Re-apply patch #1 (the
vmExtensionsfield + thevmConfig.extensions = …forward), patch #2 (LinuxProcess.killProcessGroup(_:)), and patch #3 (thesetupIOstdio-connection log + itsimport os/nucleicIOLog). Grep for[Nucleic vendored patch]to find every site. - Update the commit hash above and in the root
Package.swiftcomment. swift buildand run the balloon tests.