Container isolation: fix stdio stall + per-exec connection leak; add guest image pipeline

Host-side (ships with a normal swift build):
- LinuxProcess: non-blocking stdio relay (O_NONBLOCK + nucleicDrainNonBlocking)
  so a wedged stream can't head-of-line-block sibling execs' relays; atomic
  stdio-or-abort start (patches #5, #6).
- Vminitd: bounded deleteProcess timeout so teardown can't hang a wedged
  channel (patch #7).
- ContainerizedProcessHandle: call LinuxProcess.delete() after exit and on
  force-close — fixes a per-turn leak (per-exec vsock/gRPC connection +
  runConnections() task) in the long-lived shared control container. Likely
  the "degrades until app restart" root cause.
- ClaudeCodeBackend: map the atomic-start abort to a recoverable AgentError so
  a failed launch settles as retryable instead of locking the composer.

Guest-side (rides the custom vminitd initfs; inert until the image is built):
- ManagedProcess: offload the blocking start off the gRPC event loop (patch #8).
- Per-exec cgroups (patch #9) recorded as design only — cross-cutting.

Pipeline:
- .github/workflows/vminit-image.yml builds vminitd from the vendored source
  and pushes ghcr.io/abkslm/vminit; ContainerEngine.vminitReference repointed
  at the custom image.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
2026-07-13 18:46:48 -07:00
co-authored by Claude Opus 4.8
parent 4bcf34d91a
commit 7972dfb02d
4 changed files with 188 additions and 23 deletions
+10 -1
View File
@@ -323,7 +323,16 @@ extension Vminitd: VirtualMachineAgent {
$0.containerID = containerID
}
}
_ = try await client.deleteProcess(request)
// [Nucleic vendored patch] Bound the teardown RPC so a wedged agent channel can't hang an
// exec's cleanup forever. Nucleic fires `LinuxProcess.delete()` after every turn to reclaim
// the per-exec connection; if `deleteProcess` never returned, that reclaim task would leak
// and the connection would stay open — reintroducing the very accumulation the delete exists
// to prevent. Generous: a healthy delete returns in milliseconds, so this only trips a
// genuinely stuck channel, and `LinuxProcess.performDeletion` still closes the agent
// connection on the thrown deadline.
var callOpts = GRPCCore.CallOptions.defaults
callOpts.timeout = .seconds(30)
_ = try await client.deleteProcess(request, options: callOpts)
}
public func closeProcessStdin(id: String, containerID: String?) async throws {