Merge nucleic/olive-ember-seal-q7vk into dev
This commit is contained in:
+31
-3
@@ -6,7 +6,11 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
|
||||
|
||||
## Open
|
||||
|
||||
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale
|
||||
(none)
|
||||
|
||||
## Closed (recent)
|
||||
|
||||
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork)
|
||||
|
||||
- **Found**: narOS N0 rootfs validation (first real-world install run under the
|
||||
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
|
||||
@@ -25,8 +29,32 @@ stays open until fixed in the fork (or upstream) and re-verified by the corpus.
|
||||
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
|
||||
POSIX shells treat variable values as byte strings.
|
||||
- **Severity**: silent data corruption (not a parse failure, so the §4.1
|
||||
bash-fallback cannot catch it). Needs a corpus case (`utf8-bytes-passthru`)
|
||||
and a byte-preservation sweep of read/expansion/heredoc paths.
|
||||
bash-fallback cannot catch it).
|
||||
- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event`
|
||||
read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode —
|
||||
before pushing into the result `String` (upstream had a `TODO(utf-8)` on the
|
||||
field). Valid UTF-8 input was thus re-encoded byte-by-byte.
|
||||
- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers):
|
||||
lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are
|
||||
read (blocking, matching `read`'s wait-for-delimiter behavior; a
|
||||
non-continuation byte is pushed back), and valid sequences round-trip
|
||||
byte-identically. bash's `-n`-counts-bytes semantics are preserved for free
|
||||
because the char-limit check measures `line.len()` (UTF-8 byte length).
|
||||
Candidate for upstreaming (fixes upstream's own TODO).
|
||||
- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the
|
||||
per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes;
|
||||
byte-fidelity there would need `Vec<u8>` plumbing shell-wide). bash keeps raw
|
||||
bytes. Not corpus-gated; revisit only if it bites in practice.
|
||||
- **Verified**: repro now byte-identical to bash; split-across-writes
|
||||
continuation reassembly correct; `read -n 3` byte counting identical to bash
|
||||
(C locale); the original real-world failure (`update-ca-certificates` under
|
||||
nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including
|
||||
three new cases (`utf8-read-passthru`, `utf8-while-read-file`,
|
||||
`utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat
|
||||
suite failure set **identical to the pristine-HEAD baseline rebuilt in the
|
||||
same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing
|
||||
case that passes 10/10 standalone).
|
||||
- **Regression tests**: the three `utf8-*` corpus cases above.
|
||||
|
||||
## Closed
|
||||
|
||||
|
||||
Reference in New Issue
Block a user