Files
nash/corpus/DIVERGENCES.md
T

86 lines
5.1 KiB
Markdown
Raw Normal View History

# nash ↔ bash divergence log
Live record of behavioral divergences found by the replay harness
(`replay.py`, corpus in `corpus.jsonl`). Per docs/NASH.md §10.3/§11, each entry
stays open until fixed in the fork (or upstream) and re-verified by the corpus.
## Open
(none)
## Closed (recent)
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork)
- **Found**: narOS N0 rootfs validation (first real-world install run under the
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
(`update-ca-certificates`, a `while read`-over-conf loop) mangled the UTF-8
filename `NetLock_Arany_=Class_Gold=_Főtanúsítvány.crt` and failed the
whole image build. Worked around structurally in `os/mkimage/build-rootfs.sh`
(the divert is now always the image's last configure step), but any
*runtime* `apt install` inside narOS still runs postinsts under nash, so this
blocks narOS/M3 forcing gates until fixed.
- **Repro** (nash `0.4.0` musl arm64, empty locale):
`echo 'Főtanúsítvány' | nash -c 'while read x; do echo "$x"; done'` —
bash emits the input bytes unchanged (`F \305\221 t a n \303\272 …`); nash
emits each byte Latin-1→UTF-8 double-encoded (`F \303\205 \302\221 …`).
- **Suspected root cause**: brush decodes input bytes to `String` with a lossy/
Latin-1 assumption on the `read` path (or at word-splitting) instead of
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
POSIX shells treat variable values as byte strings.
- **Severity**: silent data corruption (not a parse failure, so the §4.1
bash-fallback cannot catch it).
- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event`
read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode —
before pushing into the result `String` (upstream had a `TODO(utf-8)` on the
field). Valid UTF-8 input was thus re-encoded byte-by-byte.
- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers):
lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are
read (blocking, matching `read`'s wait-for-delimiter behavior; a
non-continuation byte is pushed back), and valid sequences round-trip
byte-identically. bash's `-n`-counts-bytes semantics are preserved for free
because the char-limit check measures `line.len()` (UTF-8 byte length).
Candidate for upstreaming (fixes upstream's own TODO).
- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the
per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes;
byte-fidelity there would need `Vec<u8>` plumbing shell-wide). bash keeps raw
bytes. Not corpus-gated; revisit only if it bites in practice.
- **Verified**: repro now byte-identical to bash; split-across-writes
continuation reassembly correct; `read -n 3` byte counting identical to bash
(C locale); the original real-world failure (`update-ca-certificates` under
nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including
three new cases (`utf8-read-passthru`, `utf8-while-read-file`,
`utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat
suite failure set **identical to the pristine-HEAD baseline rebuilt in the
same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing
case that passes 10/10 standalone).
- **Regression tests**: the three `utf8-*` corpus cases above.
## Closed
### D1 — IFS word-splitting applied to literal words (fixed in fork)
- **Found**: M0 corpus run (`ifs-split`), brush-shell-v0.4.0.
- **Repro**: `IFS=,; set -- a,b,c; echo $1` → bash `a b c`, nash `a`.
- **Root cause**: brush applied IFS splitting to *literal* words, not just
expansion results: under `IFS=,`, `echo a,b,c` printed `a b c` (bash:
`a,b,c`) and `set -- a,b,c` received 3 args (bash: 1). POSIX/bash split only
the results of parameter/command/arithmetic expansion. Mechanically,
`brush-core/src/expansion.rs` had a two-state `ExpansionPiece`
(`Splittable`/`Unsplittable`) conflating "field-splittable" with
"glob-active", and literal text was marked `Splittable`.
- **Fix**: added a third piece state `LiteralText` (never field-split, glob
chars active) and produce it for unquoted literal text, the no-expansion-chars
fast path, and retained-backslash escapes (`// nash:` markers in
`expansion.rs`). Corpus back to 100% (94/94); candidate for upstreaming.
- **Regression tests**: corpus `ifs-split` (+ `glob-all`, `glob-txt`,
`cond-pattern`, `case` guard the glob/pattern side); brush compat suite.
- **Fix fallout (caught by the brush compat suite, both fixed)**: (a) text
substituted by `${v:-word}`/`${v:+word}`/`${v:=word}` *is* an expansion result
and must split — restored via a boundary conversion in the
`ParameterExpansion` arm; (b) `compgen -W` splits its word-list string as
data — restored via `ExpanderOptions.field_split_literal_text`. Final suite:
1684 succeeded / 6 failed, failure set identical to the pristine-brush
baseline in this container (environmental only); 3 upstream known-fail IFS
tests now pass (markers flipped in `ifs.yaml`).