Files
nucleic-purpose-classifier/datagen-prompt-2.md
T

88 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Corrective-batch generation prompt (round 2)
Round 1 (`data/purpose-prompts.jsonl`, 9,217 records) was validated with
`validate-data.py`; this prompt generates a *complementary* set that repairs the measured
deficits. The quotas below are computed so that the MERGED dataset (~12,217 records) lands
on the base prompt's aggregate targets. Because these batches are deliberately
counter-skewed, `validate-data.py` per-batch *share warnings* are expected on them;
per-batch *errors* (openers, length floors, format) must still be zero. Separately, in
curation (not generation): delete the 3 fixture-duplicate lines round 1 shipped, resolve
the 17-record batch remainder, and write the generation manifest.
---
You are generating ROUND 2 of a labeled dataset of prompts that software developers type
into a coding-agent app (like Claude Code) to START a new chat. Follow every rule in the
round-1 brief — first-messages only, the 8 labels and their boundary rules, the JSONL
schema {prompt, purpose, secondary, mixed, difficulty, slice, lang}, the anti-patterns
(no label leakage, no mid-conversation replies, no assistant-directed meta) — EXCEPT
where this brief overrides it. Round 1 drifted in measurable ways; round 2 exists to pull
the merged dataset back on target, so these overrides are hard requirements, not
preferences.
## What round 1 got wrong (so you don't repeat it)
1. It opened prompts with the same verbs constantly: "Plan…", "Implement…", "Review…",
"Explain…", "Add…", "Build…", "Audit…". 134 windows broke the opener cap.
2. It wrote almost everything at comfortable middle length — too few terse prompts, far
too few long ones, and its "pasted-context" examples were mostly under 100 tokens.
3. It over-produced backendImpl and core-slice examples; under-produced pasted-context,
mixed, boundary, refactor, review, and writing.
4. It stayed 99% English in most batches.
## Hard per-batch quotas (each batch = exactly 200 lines)
Slice counts per batch (deliberately counter-skewed; do not "fix" them toward the
round-1 targets):
- `core`: 58
- `boundary`: 52
- `pasted-context`: 50 — and every one's pasted material is genuinely 100–400 tokens
(real-looking stack traces, log tails, failing-test output, diff hunks, ticket text).
Count the tokens; round 1's 60–99-token "pastes" are rejected.
- `mixed`: 24 (`mixed: true`, secondary set — this is the 12% cap exactly)
- `vague-eval`: 16
Primary-label counts per batch (non-vague, sums to 184 — rebalances round 1's skew):
- writing: 30, review: 29, refactor: 29, quickFix: 25, frontendImpl: 22, planning: 22,
debugging: 18, backendImpl: 9
Length floors per batch (both count toward whatever slice they belong to):
- at least 60 prompts under 8 words
- at least 80 prompts over 60 estimated tokens (the 50 pasted-context entries count;
the other 30+ must come from long conversational core/boundary/mixed prompts —
rambling context-setting, multi-requirement asks, bullet-listed briefs)
Language per batch: 14–18 lines non-English, spread across es, de, fr, pt, zh, ja
(natural developer code-switching with English tech nouns; put the BCP-47 tag in `lang`).
## Opening-word discipline (this is where round 1 failed hardest)
- HARD CAP: no opening word may start more than 2 prompts in any 50 consecutive lines.
The words plan, implement, review, explain, add, build, audit, fix, create, write,
refactor, update are on a watch list — treat each as nearly exhausted before you start.
- Reach the same intents through other doors: start from the noun ("the checkout retry
logic is duplicated in three places…"), the symptom ("payments double-charge when…"),
a question ("why does the exporter…", "is there a cleaner way to…"), the artifact
("this stack trace keeps showing up:"), a stakeholder ("PM wants a writeup of…"),
lowercase mid-thought fragments ("ok so the settings pane…"), a file path
("src/sync/reconcile.ts is 900 lines now…"), or pasted material first with the ask at
the end.
- Vary sentence *shape*, not just the first token: imperative, question, complaint,
observation + ask, list of constraints, apology-then-ask, all in the mix.
## Confusion pairs to emphasize in `boundary` (round 1 under-served these)
refactor↔quickFix (renames, small restructures), review↔writing (explain vs summarize
vs document), writing↔backendImpl (docs about APIs), review↔debugging (why-questions
about behavior that may or may not be broken), planning↔refactor (restructure strategy
vs restructure execution). Keep planning↔backendImpl light — round 1 already produced
plenty.
## Process
15 batches of exactly 200 lines (3,000 total). Before each batch, silently pick 3 fresh
domains + 2 registers + 1 emphasized confusion pair; never repeat a combination. After
each batch, verify against THIS brief's quotas — slice counts, label counts, both length
floors, the 2-per-50 opener cap, non-English count — and fix violations before emitting.
JSONL only; no numbering, no fences, no commentary.