Root-cause and fix the four reported container regressions plus two
adjacent confirmed bugs.
- Memory balloon (CPU 100% + output freeze that never recovered): the
autoballoon drove the whole-VM target from a per-container cgroup figure
with no guest swap, spinning a swapless guest in perpetual direct
reclaim. Default memoryManagement to off; make the target whole-VM-aware
(reserveBytes) so it never inflates below the working set plus the
guest's non-cgroup footprint; deflate the balloon on a failed stats read
instead of freezing it inflated.
- Stop button: signal the agent's whole process group (new vendored
LinuxProcess.killProcessGroup, negative pid) so forked children die too;
replace the unbounded wait() in every teardown/shutdown with a bounded
terminate() that escalates SIGTERM -> SIGKILL; interrupt escalates to a
group kill so a wedged agent always stops.
- MCPApprovalServer port-0 race: single-flight start(host:) so concurrent
sessions sharing one control-container server all receive the real bound
port; publish listener+port only after .ready (a failed bind no longer
pins a stale port 0); guard the Claude call site against port 0.
- Container CPU metric: divide the CPU delta by the actual measured window
instead of a fixed 200 ms, so it stops over-reading under load.
- Command interceptor: drop the ~40 coreutil Node shims so cat/grep/etc
run their native binaries (no Node-per-command); the bash tracer still
records them as metadata.
- Make the bash command tracer opt-in (commandTracingEnabled, default
off) — the per-command DEBUG trap only activates when enabled; the
git/gh interception the conflict/merge system relies on stays always-on.
Co-Authored-By: Claude Opus 4.8 <[email protected]>