Merge nucleic/eager-glass-wren-mrb5 into dev

This commit is contained in:
2026-07-29 19:00:34 -07:00
commit 728766f1a5
+109
View File
@@ -0,0 +1,109 @@
# Synthetic-data generation prompt for the purpose classifier
The prompt below is fed verbatim to a frontier-model agent to produce training/eval data
per docs/PURPOSE_CLASSIFIER.md §4.1. Record the generating model, date, and batch topics
in the generation manifest alongside the output. The 82 shipped fixtures
(Tests/NucleicCoreTests/Fixtures/purpose-prompts.json) are eval-only and must NOT be
pasted into the generator's context (contamination).
---
You are generating a labeled dataset of prompts that software developers type into a
coding-agent app (like Claude Code) to start or continue a chat. Each example is ONE
prompt a real developer might send, labeled with what the prompt is FOR. The data trains
a small on-device classifier, so realism and diversity matter more than polish; label
precision matters more than anything.
## Labels (choose the primary purpose; definitions are exhaustive)
- `planning` — asking for architecture, design docs, RFCs, migration strategy,
roadmaps, breaking work into milestones. The deliverable is a PLAN or DESIGN, not code.
- `backendImpl` — implementing server/API/data/algorithm/system/CLI code: endpoints,
schemas, migrations, queues, caches, auth flows, parsers, background jobs.
- `frontendImpl` — implementing UI: views, components, styling, layout, animation,
themes, screens, visual polish. If the deliverable is something you SEE, it's frontend.
- `quickFix` — a typo, version bump, config tweak, flag flip, one-liner, or a small
contained bugfix the author already understands. Small scope, known change.
- `refactor` — restructuring without behavior change: rename/extract/split/consolidate/
dedupe/decouple/simplify. The author expects identical behavior after.
- `debugging` — diagnosing a failure the author does NOT yet understand: crashes, stack
traces, regressions, flaky tests, hangs, leaks, wrong output, "why does X happen".
- `review` — reading/judging/explaining EXISTING code or designs: code review, audits,
"what does X do", "is this safe", comparisons, walkthroughs. No code changes requested.
- `writing` — producing prose: docs, READMEs, commit messages, PR descriptions, release
notes, changelogs, summaries, translations, doc comments.
Boundary rules (apply in this order when two labels tempt you):
1. "Fix" + author already knows the change → `quickFix`. "Fix" + cause unknown /
symptoms described → `debugging`.
2. Rename/restructure "across the codebase" or preserving behavior → `refactor`, even
though a single rename in one file reads as `quickFix`.
3. Docs/comments/prose about code → `writing`, even when the subject is an API or
backend concept ("update the API docs" is `writing`).
4. "Plan/design/architect X" → `planning` even when X is backend or frontend work.
"Plan and implement X" → primary is `planning`, secondary is the implementation label.
5. A pure question about existing behavior → `review` unless something is BROKEN, then
`debugging`.
## Output format — strict JSONL, one object per line, no commentary
{"prompt": "...", "purpose": "<primary label>", "secondary": "<label or null>",
"mixed": <bool>, "difficulty": <0.0-1.0>, "slice": "<slice tag>", "lang": "<bcp47>"}
- `secondary`/`mixed`: only when the prompt genuinely asks for two purposes ("plan and
implement…", "fix the crash and write a regression test note"). At most ~12% of
examples; `mixed` false → `secondary` null.
- `difficulty`: how much model capability the TASK described would need. Rubric:
0.0–0.2 trivial (typo, one-liner); 0.3–0.5 routine scoped work; 0.6–0.8 multi-file /
multi-constraint / gnarly diagnosis; 0.9–1.0 long-horizon, architectural, high-risk.
Grade the task, not the prompt's length.
- `slice`, one of:
- `core` — clear, single-purpose prompts (≈55% of output)
- `boundary` — deliberately near a label boundary per the rules above (≈20%)
- `mixed` — genuine two-purpose prompts (≈10%)
- `pasted-context` — the ask is buried in pasted material: stack traces, log tail,
diff hunk, failing test output, a TODO list. Reproduce the pasted material
realistically, 100–400 tokens (≈10%)
- `vague-eval` — terse/ambiguous prompts with no recoverable purpose ("continue",
"make it pop", "do the thing we discussed"). Label with your best guess anyway;
these are held out of training by the pipeline (≈5%)
## Diversity requirements (enforced per batch)
- Length: from 2 words to ~400 tokens; at least 15% under 8 words, at least 15% over
60 tokens.
- Register: terse imperatives, polite asks, stream-of-consciousness, bullet lists,
mid-thought fragments, sloppy typing with real typos (don't correct them), pasted
Slack/ticket text. NEVER the same opening verb more than 3 times per 50 examples.
- Domain: rotate web frontend, iOS/macOS (SwiftUI), Android, backend (Go/Rust/Python/
TS/Java), data/ML, infra/DevOps, embedded, games, databases, CLI tools.
- Tech nouns: use real, current, varied technology names and file paths; invent
plausible project-specific names (components, services, feature flags) so examples
aren't keyword-matchable.
- Language: ~95% English; ~5% spread across es, de, fr, pt, zh, ja (natural developer
usage, often code-switched with English tech nouns).
- Confusion pairs to mine deliberately in the `boundary` slice, several dozen each:
writing↔backendImpl (docs about APIs), quickFix↔refactor (renames), review↔debugging
(why-questions), planning↔backendImpl ("plan and implement"), frontendImpl↔quickFix
(small UI tweaks), review↔writing (explain vs summarize).
## Anti-patterns (rejected in QA)
- Label leakage: prompts must never contain the label word used AS a label hint
("refactor this" is fine and common; "this is a refactor task:" is not).
- Template smell: recycled sentence skeletons with one noun swapped; enumerated
"Task 47:" prefixes; uniform lengths; every example ending in a period.
- Impossible labels: prompts a human labeler couldn't defend from the text alone
(except the `vague-eval` slice, where that's the point).
- Assistant-directed meta ("classify this prompt", "as an AI") — these are prompts TO a
coding agent, about code.
## Process
Produce the dataset in batches of 200 lines. Before each batch, silently pick a fresh
combination of 3 domains + 2 registers + 1 confusion pair to emphasize, so no two
batches have the same texture (do not print your picks — JSONL lines only). After each
batch, self-check against the anti-patterns and the slice/label distributions, and fix
violations before emitting. Across the full run, keep primary labels within ±15% of
uniform across the 8 classes (the vague-eval slice is exempt). Target total: 8,000
lines. Do not number examples. Do not wrap output in markdown fences. JSONL only.