Per-exec cgroups (guest patch #9): scope a session's OOM/CPU/fork-bomb to itself

Restructures the guest cgroup layout so each exec gets its OWN child cgroup
(/container/<id>/<execID>) with memory.oom.group=1, a fair cpu.weight, and a
pids.max backstop — so one control session can't OOM-kill, starve, or fork-bomb
its siblings in the shared container. The container init moves to its own leaf
so the container cgroup can delegate controllers to children (cgroup v2
no-internal-process rule). New Cgroup2Manager helpers: setOomGroup/setCpuWeight/
setPidsMax/remove.

Best-effort with graceful fallback: any failure in the per-exec setup wipes the
partial state and reverts to today's flat layout, and each exec falls back to the
container cgroup — a cgroup hiccup degrades to current behavior, never a failed
start.

COMPILE-VERIFIED via the musl cross-build; NOT yet runtime-validated. Built as
image tag -nucleic2; vminitReference stays on the validated -nucleic1 until
-nucleic2 is checked in a real container. A hard host-configured per-exec
memory.max (exec-RPC resources field) remains a follow-up.

Co-Authored-By: Claude Opus 4.8 <[email protected]>
This commit is contained in:
2026-07-13 20:03:20 -07:00
co-authored by Claude Opus 4.8
parent 0f9957d5e5
commit 2eb563c90c
4 changed files with 130 additions and 32 deletions
@@ -267,6 +267,33 @@ public struct Cgroup2Manager: Sendable {
fileName: "memory.low")
}
/// [Nucleic vendored patch] Make the kernel OOM-killer treat this cgroup as an atomic unit: when a
/// memory limit (this cgroup's or an ancestor's) forces an OOM, the whole cgroup's process tree is
/// killed together rather than one victim. Used to scope a runaway exec's OOM to that exec so
/// sibling execs in the same container survive.
package func setOomGroup(_ enabled: Bool) throws {
try Self.writeValue(path: self.path, value: enabled ? "1" : "0", fileName: "memory.oom.group")
}
/// [Nucleic vendored patch] Relative CPU share under contention (cgroup v2 `cpu.weight`, 1…10000,
/// default 100). Equal weights give each exec a fair slice so one busy session can't starve
/// siblings of CPU.
package func setCpuWeight(_ weight: UInt64) throws {
try Self.writeValue(path: self.path, value: String(weight), fileName: "cpu.weight")
}
/// [Nucleic vendored patch] Cap the pids in this cgroup (`pids.max`) — a fork-bomb backstop so one
/// exec can't exhaust the pid space and wedge its siblings.
package func setPidsMax(_ max: UInt64) throws {
try Self.writeValue(path: self.path, value: String(max), fileName: "pids.max")
}
/// [Nucleic vendored patch] Remove this cgroup directory (rmdir). The cgroup must already be empty
/// of processes and child cgroups. Best-effort partial-setup cleanup for the per-exec layout.
package func remove() throws {
try FileManager.default.removeItem(at: self.path)
}
package func getMemoryEvents() throws -> MemoryEvents {
let content = try readFileContent(fileName: "memory.events")
let values = parseKeyValuePairs(content)