Merge nucleic/upbeat-opal-viper-cc8l into main
build / build (push) Canceled after 0s

This commit is contained in:
2026-08-07 05:11:01 -07:00
parent ee19744496
commit 6b0c01b74b
8 changed files with 395 additions and 28 deletions
+12 -2
View File
@@ -284,8 +284,18 @@ scheduler treats as transient back-pressure rather than a failure.
### Timeouts
* A slot in `.provisioning(since:)` longer than `bootTimeoutSeconds` (default
300) is torn down. Covers a guest that never gets a lease, never starts `sshd`,
or hangs in Setup Assistant.
900) is torn down. Covers a guest that never gets a lease, never starts `sshd`,
or hangs in Setup Assistant. The default is deliberately generous: several
Virtualization guests sharing one host push a boot from tens of seconds into
minutes, and a limit below the worst case does not time out a bad boot, it
livelocks — each replacement clone starts from zero and adds load, so the next
boot is slower still and no runner ever registers.
* The lifecycle's own `waitForLease` and `waitForSSH` budgets are derived from
what is *left* of that deadline, not from a fresh copy of it. Given the full
`bootTimeoutSeconds` their deadlines would fall after the planner's, so the
planner would always cancel first and the specific error — which host, how
many attempts, what the last one said — would be discarded in favour of a bare
cancellation.
* A slot in `.running(jobHint:since:)` longer than `jobTimeoutMinutes` (default
120) is torn down. Covers a job that hangs. This is comfortably below Gitea's
own `ABANDONED_JOB_TIMEOUT` (24 h), so our teardown always happens first and
+1 -1
View File
@@ -277,7 +277,7 @@ is required; the file form wins over the inline form when both are present.
| `pollIntervalSeconds` | `5` | How often to poll the queued-jobs API. |
| `reconcileIntervalSeconds` | `300` | How often to sweep Gitea for orphaned runner registrations from uncleanly-killed VMs. |
| `jobTimeoutMinutes` | `120` | Wall-clock limit for one job; the VM is destroyed when exceeded. |
| `bootTimeoutSeconds` | `300` | Time allowed from VM start to a usable SSH connection. |
| `bootTimeoutSeconds` | `900` | Time allowed from VM start to a usable SSH connection. |
**`guest`**
+58
View File
@@ -26,6 +26,7 @@ gitea-macos-runner service status
| VM boots but never gets an IP | DHCP lease not yet written, or Local Network privacy denial (macOS 15+) | Check `/var/db/dhcpd_leases`; grant Local Network permission or pre-authorize the subnet |
| Runner not listed under Privacy & Security → Local Network | Expected — the list is populated only after the app first attempts a local connection; it cannot be pre-approved | Boot one VM by hand from a GUI Terminal to create the entry, or (better on CI) allowlist the subnet with `defaults write com.apple.network.local-network` |
| SSH times out on a freshly built image | Guest macOS < 27, so provisioning options were ignored and Setup Assistant is waiting | Rebuild the image from a macOS **27+** IPSW |
| VMs boot in a loop; every teardown says `reason=cancelled` and nothing is logged between the lease and the teardown | `scheduler.bootTimeoutSeconds` is below the guest's *worst-case* boot on a contended host, so each clone is killed while still starting — and each replacement makes the next one slower | Raise `scheduler.bootTimeoutSeconds` (default 900) and reduce the number of concurrent guests; see [The daemon boots VMs forever](#the-daemon-boots-vms-forever-and-every-teardown-says-reasoncancelled) |
| `ssh failed: cannot connect … No route to host) (errno: 65)` part-way through provisioning | macOS 15+ Local Network privacy blocking the app — the grant is keyed on the executable's UUID, so `make install` withdraws it | Allowlist the subnet (`192.168.64.0/18`) and **reboot**; see [SSH fails with "No route to host" mid-run](#ssh-fails-with-no-route-to-host-errno-65-mid-run) |
| Allowlist is set but guests are still unreachable | It names `192.168.64.0/24` while the NAT has moved to `192.168.65.x` | Widen it to `192.168.64.0/18` and reboot; `doctor` now warns about too-narrow allowlists |
| `SecKeyCreateRandomKey` / "Interaction is not allowed" | `login.keychain` is locked — no GUI session | Run as a LaunchAgent in an unlocked GUI session; enable auto-login |
@@ -195,6 +196,63 @@ with this builder.
---
## The daemon boots VMs forever and every teardown says `reason=cancelled`
**Symptom.** A job is queued, the daemon is running, and the log repeats the same three lines with a
new runner name each time — but no runner ever appears in Gitea:
```
info orchestrator: job=1 runner=macos-vm-1cd8e83f… slot=0 booting VM
info orchestrator: ip=192.168.65.233 slot=0 guest leased address
info orchestrator: reason=cancelled slot=0 tearing down slot
```
Note what is missing: nothing between the lease and the teardown, and a teardown reason that names
no cause.
**Cause.** The guest takes longer to reach `sshd` than `scheduler.bootTimeoutSeconds` allows, so the
scheduler tears the slot down while it is still coming up — usually seconds before it would have
succeeded. This is not a timeout that fires once; it is a **livelock**. The replacement clone starts
from zero *and* adds load to an already contended host, so the next boot is slower still and the
loop never converges.
Several Virtualization guests on one Mac is enough to cause it: a guest that reaches SSH in 40
seconds on an idle host can take four or five minutes when it is sharing the machine, and each slot
holds 4 vCPU and 8 GB for the whole attempt. Check with `uptime` inside a guest — a load average in
the tens means the guest is starved, not broken.
**Fix.**
1. Raise `scheduler.bootTimeoutSeconds`. The default is 900; treat it as a ceiling on the guest's
*worst* case, not its typical one. Timing out too early costs far more than noticing a genuinely
wedged guest late.
2. Reduce contention. Count what is actually running:
```sh
ps -Ao pid,rss,etime,comm | grep -i -e virtual -e vmnet
```
Virtualization guests belonging to *other* tools compete for the same cores and the same two-VM
macOS limit. Shut down what you are not using, or lower `scheduler.maxConcurrentVMs`.
3. Confirm the guest itself is fine, independently of the daemon, with `doctor` — its `guest ssh`
check authenticates against whichever slot currently holds a lease:
```sh
gitea-macos-runner doctor
```
**If you are on an older build**, upgrade: the empty gap in that log was three bugs, all now fixed.
`waitForSSH` was handed the full `bootTimeoutSeconds` even though the scheduler's clock had started
before the clone — so the scheduler always fired first and `waitForSSH`'s own error was unreachable;
its per-attempt failures were only ever reported in that unreachable error; and the teardown
overwrote the planner's reason with `cancelled`. Current builds log `waiting for guest ssh` with the
attempt count and the last error while it is happening, and report
`reason="boot timeout: provisioning for 312s (limit 300s)"`.
---
## `SecKeyCreateRandomKey` / "Interaction is not allowed"
**Symptom.** The daemon starts but fails during VM setup with a Security-framework error mentioning