Files
nash/corpus/DIVERGENCES.md
T

86 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# nash ↔ bash divergence log
Live record of behavioral divergences found by the replay harness
(`replay.py`, corpus in `corpus.jsonl`). Per docs/NASH.md §10.3/§11, each entry
stays open until fixed in the fork (or upstream) and re-verified by the corpus.
## Open
(none)
## Closed (recent)
### D2 — non-ASCII bytes re-encoded through `read`/`echo` under C/empty locale (fixed in fork)
- **Found**: narOS N0 rootfs validation (first real-world install run under the
divert): with `/bin/sh → nash`, `ca-certificates`' postinst
(`update-ca-certificates`, a `while read`-over-conf loop) mangled the UTF-8
filename `NetLock_Arany_=Class_Gold=_Főtanúsítvány.crt` and failed the
whole image build. Worked around structurally in `os/mkimage/build-rootfs.sh`
(the divert is now always the image's last configure step), but any
*runtime* `apt install` inside narOS still runs postinsts under nash, so this
blocks narOS/M3 forcing gates until fixed.
- **Repro** (nash `0.4.0` musl arm64, empty locale):
`echo 'Főtanúsítvány' | nash -c 'while read x; do echo "$x"; done'` —
bash emits the input bytes unchanged (`F \305\221 t a n \303\272 …`); nash
emits each byte Latin-1→UTF-8 double-encoded (`F \303\205 \302\221 …`).
- **Suspected root cause**: brush decodes input bytes to `String` with a lossy/
Latin-1 assumption on the `read` path (or at word-splitting) instead of
keeping raw bytes; on output the char sequence is re-encoded as UTF-8.
POSIX shells treat variable values as byte strings.
- **Severity**: silent data corruption (not a parse failure, so the §4.1
bash-fallback cannot catch it).
- **Root cause (confirmed)**: `brush-builtins/src/read.rs` `InputReader::read_event`
read one byte at a time and did `self.buffer[0] as char` — a Latin-1 decode —
before pushing into the result `String` (upstream had a `TODO(utf-8)` on the
field). Valid UTF-8 input was thus re-encoded byte-by-byte.
- **Fix**: incremental UTF-8 decoding in `read_event` (`// nash(D2)` markers):
lead byte classifies the sequence length (0xC2–0xF4), continuation bytes are
read (blocking, matching `read`'s wait-for-delimiter behavior; a
non-continuation byte is pushed back), and valid sequences round-trip
byte-identically. bash's `-n`-counts-bytes semantics are preserved for free
because the char-limit check measures `line.len()` (UTF-8 byte length).
Candidate for upstreaming (fixes upstream's own TODO).
- **Residual (accepted)**: *invalid* UTF-8 input bytes still degrade to the
per-byte Latin-1 mapping (a Rust `String` cannot hold raw invalid bytes;
byte-fidelity there would need `Vec<u8>` plumbing shell-wide). bash keeps raw
bytes. Not corpus-gated; revisit only if it bites in practice.
- **Verified**: repro now byte-identical to bash; split-across-writes
continuation reassembly correct; `read -n 3` byte counting identical to bash
(C locale); the original real-world failure (`update-ca-certificates` under
nash) exits 0 with all 141 certs processed; corpus 97/97 = 100% including
three new cases (`utf8-read-passthru`, `utf8-while-read-file`,
`utf8-cmdsub-roundtrip`); brush-builtins read tests 20/20; brush compat
suite failure set **identical to the pristine-HEAD baseline rebuilt in the
same container** (29 shared environmental failures, ±1 flaky SIGPIPE-timing
case that passes 10/10 standalone).
- **Regression tests**: the three `utf8-*` corpus cases above.
## Closed
### D1 — IFS word-splitting applied to literal words (fixed in fork)
- **Found**: M0 corpus run (`ifs-split`), brush-shell-v0.4.0.
- **Repro**: `IFS=,; set -- a,b,c; echo $1` → bash `a b c`, nash `a`.
- **Root cause**: brush applied IFS splitting to *literal* words, not just
expansion results: under `IFS=,`, `echo a,b,c` printed `a b c` (bash:
`a,b,c`) and `set -- a,b,c` received 3 args (bash: 1). POSIX/bash split only
the results of parameter/command/arithmetic expansion. Mechanically,
`brush-core/src/expansion.rs` had a two-state `ExpansionPiece`
(`Splittable`/`Unsplittable`) conflating "field-splittable" with
"glob-active", and literal text was marked `Splittable`.
- **Fix**: added a third piece state `LiteralText` (never field-split, glob
chars active) and produce it for unquoted literal text, the no-expansion-chars
fast path, and retained-backslash escapes (`// nash:` markers in
`expansion.rs`). Corpus back to 100% (94/94); candidate for upstreaming.
- **Regression tests**: corpus `ifs-split` (+ `glob-all`, `glob-txt`,
`cond-pattern`, `case` guard the glob/pattern side); brush compat suite.
- **Fix fallout (caught by the brush compat suite, both fixed)**: (a) text
substituted by `${v:-word}`/`${v:+word}`/`${v:=word}` *is* an expansion result
and must split — restored via a boundary conversion in the
`ParameterExpansion` arm; (b) `compgen -W` splits its word-list string as
data — restored via `ExpanderOptions.field_split_literal_text`. Final suite:
1684 succeeded / 6 failed, failure set identical to the pristine-brush
baseline in this container (environmental only); 3 upstream known-fail IFS
tests now pass (markers flipped in `ifs.yaml`).