Merge nucleic/tidy-north-gecko-6wqv into dev
This commit is contained in:
+47
-1
@@ -78,6 +78,28 @@ in-tree means the patch can't be lost to a dependency re-resolve.
|
||||
the exec exits / on force-close) — without that caller, upstream never deletes execs at all and
|
||||
the shared control container leaks a connection + `runConnections()` task per turn.
|
||||
|
||||
10. **`Sources/ContainerizationOS/Socket/Socket.swift` — `acceptStream` survives transient accept
|
||||
errors.** Upstream cancelled the accept `DispatchSource` on ANY `accept(2)` failure, permanently
|
||||
ending accepting while the socket stayed **bound and listening** — a silent black hole: every
|
||||
later client `connect(2)` SUCCEEDED into the kernel backlog and hung forever unanswered. When
|
||||
the socket is the relayed control plane of a shared container, that is the "agent produced no
|
||||
output within 60s / stdio transport stalled" all-sessions wedge, fixable only by recreating the
|
||||
VM (an app restart). Transient errors — `ECONNABORTED`/`ECONNRESET` (a queued connection dying
|
||||
before accept, routine under connection churn), `EMFILE`/`ENFILE`/`ENOBUFS`/`ENOMEM` (resource
|
||||
pressure), `EINTR`/`EAGAIN` — now skip that one accept and keep listening (new
|
||||
`isTransientAcceptError`). This is SHARED code: the fix reaches the host by a normal build and
|
||||
the guest via the initfs rebuild. Marked `[Nucleic vendored patch]`.
|
||||
|
||||
11. **`Sources/Containerization/UnixSocketRelay.swift` — relay loops contain per-connection
|
||||
failures.** Both accept loops (`setupHostVsockListener` — the host half of a container's relayed
|
||||
control socket — and `setupHostVsockDial`) used to let ONE thrown per-connection dial/connect
|
||||
(host server rebinding, backlog momentarily full → `ECONNREFUSED`, fd pressure) propagate out of
|
||||
the loop, whose teardown then removed the vsock listener — permanently severing every session in
|
||||
the container from the host control plane. Each connection is now handled on its own task with
|
||||
its error logged and contained (mirroring the guest `VsockProxy`), and the pre-relay failure
|
||||
paths close both ends so a failed connection fails FAST for the peer and leaks no fds. Marked
|
||||
`[Nucleic vendored patch]`.
|
||||
|
||||
### GUEST-side patches (require rebuilding the initfs — see below)
|
||||
|
||||
Patches #1–#7 are host-side (the `Containerization` library), shipped by a normal `swift build`.
|
||||
@@ -132,6 +154,27 @@ rebuild whenever a guest patch changes. Built locally, not in CI: the host frame
|
||||
`linux.resources` at `linux.cgroupsPath`, which the patch repoints (init leaf) and clears
|
||||
accordingly.
|
||||
|
||||
12. **`vminitd/Sources/VminitdCore/VsockProxy.swift` — leak-proof, crash-proof relay connections;
|
||||
no black-hole listener.** Four fixes to the guest half of the relayed control socket (the path
|
||||
every session's MCP/approval traffic crosses in a shared container):
|
||||
- **fd leak (the root of the recurring all-sessions stall):** `cleanup` ran its two epoll
|
||||
unregisters and two `close(2)`s in one `do/catch`, so a thrown unregister SKIPPED the closes —
|
||||
leaking both connection fds. Control-plane traffic is connection-churny by design (an SSE
|
||||
`tools/call` closes its connection every gated call; every intercepted git/gh/command event is
|
||||
a short-lived connection), so the leaks accumulated until vminitd hit `EMFILE`, its accept
|
||||
path began failing, and — before patch #10 — the accept stream died with the guest socket
|
||||
still bound: every session in the container then stalled ("produced no output within 60s")
|
||||
until the VM was recreated. Each cleanup step now runs independently.
|
||||
- **double-resume crash:** both fds' epoll handlers can reach the cleanup condition; a second
|
||||
entry would resume the `CheckedContinuation` twice — a fatal trap in the VM's PID-1 agent.
|
||||
`cleanup` is now once-guarded.
|
||||
- **`try!` registrations:** an `epoll_ctl` failure crashed vminitd outright; registration
|
||||
failures now fail only that connection, releasing whatever was already set up.
|
||||
- **no black-hole listener:** if the accept loop ever ends unexpectedly, the proxy now closes
|
||||
its listener (new `listenerLoopEnded`), so peers get fail-fast refusals instead of connecting
|
||||
into a never-accepted backlog. A failed pre-relay connection is also closed explicitly.
|
||||
Marked `[Nucleic vendored patch]`.
|
||||
|
||||
## Re-vendoring a newer upstream commit
|
||||
|
||||
1. `git clone` upstream (or copy `.build/checkouts/containerization` after bumping the URL pin
|
||||
@@ -146,7 +189,10 @@ rebuild whenever a guest patch changes. Built locally, not in CI: the host frame
|
||||
stdio-or-abort guard in `start()`), patch #7 (the bounded `deleteProcess` timeout in
|
||||
`Vminitd.swift`), and patch #8 (the `ManagedProcess.start` event-loop offload in `vminitd/`). Grep
|
||||
for `[Nucleic vendored patch]` to find every site, and patch #9 (per-exec cgroups) across
|
||||
`Cgroup2Manager.swift` / `ManagedContainer.swift` / `ManagedProcess.swift`. After re-applying any
|
||||
`Cgroup2Manager.swift` / `ManagedContainer.swift` / `ManagedProcess.swift`, patch #10
|
||||
(`Socket.acceptStream` transient-error tolerance + `isTransientAcceptError`), patch #11 (the
|
||||
`UnixSocketRelay` per-connection containment + fail-fast closes), and patch #12 (the `VsockProxy`
|
||||
cleanup/`try!`/listener hardening in `vminitd/`). After re-applying any
|
||||
`vminitd/` patch, rebuild + publish the custom init image
|
||||
with `make vminit-image` + `make vminit-image-push`, and bump `ContainerEngine.vminitReference`.
|
||||
5. Update the commit hash above and in the root `Package.swift` comment.
|
||||
|
||||
Reference in New Issue
Block a user