Fix Nucleic Control container breakage: balloon, stop, interceptor

Root-cause and fix the four reported container regressions plus two
adjacent confirmed bugs.

- Memory balloon (CPU 100% + output freeze that never recovered): the
  autoballoon drove the whole-VM target from a per-container cgroup figure
  with no guest swap, spinning a swapless guest in perpetual direct
  reclaim. Default memoryManagement to off; make the target whole-VM-aware
  (reserveBytes) so it never inflates below the working set plus the
  guest's non-cgroup footprint; deflate the balloon on a failed stats read
  instead of freezing it inflated.
- Stop button: signal the agent's whole process group (new vendored
  LinuxProcess.killProcessGroup, negative pid) so forked children die too;
  replace the unbounded wait() in every teardown/shutdown with a bounded
  terminate() that escalates SIGTERM -> SIGKILL; interrupt escalates to a
  group kill so a wedged agent always stops.
- MCPApprovalServer port-0 race: single-flight start(host:) so concurrent
  sessions sharing one control-container server all receive the real bound
  port; publish listener+port only after .ready (a failed bind no longer
  pins a stale port 0); guard the Claude call site against port 0.
- Container CPU metric: divide the CPU delta by the actual measured window
  instead of a fixed 200 ms, so it stops over-reading under load.
- Command interceptor: drop the ~40 coreutil Node shims so cat/grep/etc
  run their native binaries (no Node-per-command); the bash tracer still
  records them as metadata.
- Make the bash command tracer opt-in (commandTracingEnabled, default
  off) — the per-command DEBUG trap only activates when enabled; the
  git/gh interception the conflict/merge system relies on stays always-on.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
Nucleic
2026-06-25 13:55:31 -07:00
co-authored by Claude Opus 4.8
parent 11b9825e09
commit 050605b13e
2 changed files with 32 additions and 2 deletions
+11 -2
View File
@@ -21,7 +21,15 @@ in-tree means the patch can't be lost to a dependency re-resolve.
target at runtime for automatic VM memory reclamation — see `MemoryBalloon.swift` /
`ContainerEngine` in NucleicCore.
2. **Trimmed for footprint (no behavior change).** `Tests/`, `docs/`, `examples/`, and `images/`
2. **`Sources/Containerization/LinuxProcess.swift` — process-group kill.**
`LinuxProcess` gains `killProcessGroup(_:)`, which signals the negative pid (`-pid`) so the
guest's `kill(2)` targets the exec'd process's whole **process group**, not just the leader.
Every exec is `setsid()`'d by `vmexec`, so the process is its own group leader (pgid == pid) and
a group signal reaches the children it forked. Upstream only exposes the leader-only `kill(_:)`,
which let a forked child survive a Stop in a long-lived shared container. Marked with
`[Nucleic vendored patch]`; used by `ContainerizedProcessHandle.sendSignal` in NucleicCore.
3. **Trimmed for footprint (no behavior change).** `Tests/`, `docs/`, `examples/`, and `images/`
were dropped, and the corresponding `.testTarget(...)` entries removed from `Package.swift`. The
library/executable targets we build are untouched.
@@ -32,6 +40,7 @@ in-tree means the patch can't be lost to a dependency re-resolve.
2. `rsync -a --exclude=.git --exclude=.build --exclude=.swiftpm --exclude=Tests/ --exclude=docs/ \
--exclude=examples/ --exclude=images/ <upstream>/ third_party/containerization/`
3. Remove the `.testTarget(...)` blocks from `third_party/containerization/Package.swift`.
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward).
4. Re-apply patch #1 (the `vmExtensions` field + the `vmConfig.extensions = …` forward) and patch
#2 (`LinuxProcess.killProcessGroup(_:)`). Grep for `[Nucleic vendored patch]` to find every site.
5. Update the commit hash above and in the root `Package.swift` comment.
6. `swift build` and run the balloon tests.
@@ -320,6 +320,27 @@ extension LinuxProcess {
}
}
/// [Nucleic vendored patch] Deliver a signal to the whole process GROUP led by this exec'd
/// process, not just the leader. `vmexec` `setsid()`s every exec, so the process is its own
/// session/group leader and its pgid equals its pid; a negative pid makes the guest's `kill(2)`
/// target the entire group, reaching any children the agent forked (model/turn subprocesses,
/// tool shells). `kill(_:)` above signals only the leader, so a wedged child can survive a Stop
/// in a long-lived shared container — this is the group-wide counterpart. Best-effort and
/// guarded against pid ≤ 1 (a non-positive pid would target the caller's group / every process).
public func killProcessGroup(_ signal: Signal) async throws {
let leader = self.pid
guard leader > 1 else { return }
do {
_ = try await agent.kill(pid: -leader, signal: signal.rawValue)
} catch {
throw ContainerizationError(
.internalError,
message: "failed to kill process group",
cause: error
)
}
}
/// Resize the processes pty (if requested).
public func resize(to: Terminal.Size) async throws {
do {