2026-07-29 19:00:34 -07:00
|
|
|
|
# Synthetic-data generation prompt for the purpose classifier
|
|
|
|
|
|
|
|
|
|
|
|
The prompt below is fed verbatim to a frontier-model agent to produce training/eval data
|
|
|
|
|
|
per docs/PURPOSE_CLASSIFIER.md §4.1. Record the generating model, date, and batch topics
|
|
|
|
|
|
in the generation manifest alongside the output. The 82 shipped fixtures
|
|
|
|
|
|
(Tests/NucleicCoreTests/Fixtures/purpose-prompts.json) are eval-only and must NOT be
|
|
|
|
|
|
pasted into the generator's context (contamination).
|
|
|
|
|
|
|
|
|
|
|
|
---
|
|
|
|
|
|
|
|
|
|
|
|
You are generating a labeled dataset of prompts that software developers type into a
|
2026-07-29 21:01:20 -07:00
|
|
|
|
coding-agent app (like Claude Code) to START a new chat. Each example is the OPENING
|
|
|
|
|
|
message of a fresh chat — never a reply inside an ongoing conversation — labeled with
|
|
|
|
|
|
what the prompt is FOR. This matters: the classifier runs exactly once, on the first
|
|
|
|
|
|
message, to pick the chat's model (which then stays fixed for the chat's whole life), so
|
|
|
|
|
|
the training distribution must be first-messages only. A first message may still
|
|
|
|
|
|
reference prior work the way developers really do ("continuing from yesterday's auth
|
|
|
|
|
|
refactor, …", "picking up the payment-flow bug again"), because people routinely start
|
|
|
|
|
|
fresh chats mid-project. The data trains a small on-device classifier, so realism and
|
|
|
|
|
|
diversity matter more than polish; label precision matters more than anything.
|
2026-07-29 19:00:34 -07:00
|
|
|
|
|
|
|
|
|
|
## Labels (choose the primary purpose; definitions are exhaustive)
|
|
|
|
|
|
|
|
|
|
|
|
- `planning` — asking for architecture, design docs, RFCs, migration strategy,
|
|
|
|
|
|
roadmaps, breaking work into milestones. The deliverable is a PLAN or DESIGN, not code.
|
|
|
|
|
|
- `backendImpl` — implementing server/API/data/algorithm/system/CLI code: endpoints,
|
|
|
|
|
|
schemas, migrations, queues, caches, auth flows, parsers, background jobs.
|
|
|
|
|
|
- `frontendImpl` — implementing UI: views, components, styling, layout, animation,
|
|
|
|
|
|
themes, screens, visual polish. If the deliverable is something you SEE, it's frontend.
|
|
|
|
|
|
- `quickFix` — a typo, version bump, config tweak, flag flip, one-liner, or a small
|
|
|
|
|
|
contained bugfix the author already understands. Small scope, known change.
|
|
|
|
|
|
- `refactor` — restructuring without behavior change: rename/extract/split/consolidate/
|
|
|
|
|
|
dedupe/decouple/simplify. The author expects identical behavior after.
|
|
|
|
|
|
- `debugging` — diagnosing a failure the author does NOT yet understand: crashes, stack
|
|
|
|
|
|
traces, regressions, flaky tests, hangs, leaks, wrong output, "why does X happen".
|
|
|
|
|
|
- `review` — reading/judging/explaining EXISTING code or designs: code review, audits,
|
|
|
|
|
|
"what does X do", "is this safe", comparisons, walkthroughs. No code changes requested.
|
|
|
|
|
|
- `writing` — producing prose: docs, READMEs, commit messages, PR descriptions, release
|
|
|
|
|
|
notes, changelogs, summaries, translations, doc comments.
|
|
|
|
|
|
|
|
|
|
|
|
Boundary rules (apply in this order when two labels tempt you):
|
|
|
|
|
|
1. "Fix" + author already knows the change → `quickFix`. "Fix" + cause unknown /
|
|
|
|
|
|
symptoms described → `debugging`.
|
|
|
|
|
|
2. Rename/restructure "across the codebase" or preserving behavior → `refactor`, even
|
|
|
|
|
|
though a single rename in one file reads as `quickFix`.
|
|
|
|
|
|
3. Docs/comments/prose about code → `writing`, even when the subject is an API or
|
|
|
|
|
|
backend concept ("update the API docs" is `writing`).
|
|
|
|
|
|
4. "Plan/design/architect X" → `planning` even when X is backend or frontend work.
|
|
|
|
|
|
"Plan and implement X" → primary is `planning`, secondary is the implementation label.
|
|
|
|
|
|
5. A pure question about existing behavior → `review` unless something is BROKEN, then
|
|
|
|
|
|
`debugging`.
|
|
|
|
|
|
|
|
|
|
|
|
## Output format — strict JSONL, one object per line, no commentary
|
|
|
|
|
|
|
|
|
|
|
|
{"prompt": "...", "purpose": "<primary label>", "secondary": "<label or null>",
|
|
|
|
|
|
"mixed": <bool>, "difficulty": <0.0-1.0>, "slice": "<slice tag>", "lang": "<bcp47>"}
|
|
|
|
|
|
|
|
|
|
|
|
- `secondary`/`mixed`: only when the prompt genuinely asks for two purposes ("plan and
|
|
|
|
|
|
implement…", "fix the crash and write a regression test note"). At most ~12% of
|
|
|
|
|
|
examples; `mixed` false → `secondary` null.
|
|
|
|
|
|
- `difficulty`: how much model capability the TASK described would need. Rubric:
|
|
|
|
|
|
0.0–0.2 trivial (typo, one-liner); 0.3–0.5 routine scoped work; 0.6–0.8 multi-file /
|
|
|
|
|
|
multi-constraint / gnarly diagnosis; 0.9–1.0 long-horizon, architectural, high-risk.
|
|
|
|
|
|
Grade the task, not the prompt's length.
|
|
|
|
|
|
- `slice`, one of:
|
|
|
|
|
|
- `core` — clear, single-purpose prompts (≈55% of output)
|
|
|
|
|
|
- `boundary` — deliberately near a label boundary per the rules above (≈20%)
|
|
|
|
|
|
- `mixed` — genuine two-purpose prompts (≈10%)
|
|
|
|
|
|
- `pasted-context` — the ask is buried in pasted material: stack traces, log tail,
|
|
|
|
|
|
diff hunk, failing test output, a TODO list. Reproduce the pasted material
|
|
|
|
|
|
realistically, 100–400 tokens (≈10%)
|
2026-07-29 21:01:20 -07:00
|
|
|
|
- `vague-eval` — terse/ambiguous prompts with no recoverable purpose, typed as the
|
|
|
|
|
|
first message of a new chat ("continue", "make it pop", "do the thing we
|
|
|
|
|
|
discussed" — real behavior when someone reopens work in a fresh chat). Label with
|
|
|
|
|
|
your best guess anyway; these are held out of training by the pipeline (≈5%)
|
2026-07-29 19:00:34 -07:00
|
|
|
|
|
|
|
|
|
|
## Diversity requirements (enforced per batch)
|
|
|
|
|
|
|
|
|
|
|
|
- Length: from 2 words to ~400 tokens; at least 15% under 8 words, at least 15% over
|
|
|
|
|
|
60 tokens.
|
|
|
|
|
|
- Register: terse imperatives, polite asks, stream-of-consciousness, bullet lists,
|
|
|
|
|
|
mid-thought fragments, sloppy typing with real typos (don't correct them), pasted
|
|
|
|
|
|
Slack/ticket text. NEVER the same opening verb more than 3 times per 50 examples.
|
|
|
|
|
|
- Domain: rotate web frontend, iOS/macOS (SwiftUI), Android, backend (Go/Rust/Python/
|
|
|
|
|
|
TS/Java), data/ML, infra/DevOps, embedded, games, databases, CLI tools.
|
|
|
|
|
|
- Tech nouns: use real, current, varied technology names and file paths; invent
|
|
|
|
|
|
plausible project-specific names (components, services, feature flags) so examples
|
|
|
|
|
|
aren't keyword-matchable.
|
|
|
|
|
|
- Language: ~95% English; ~5% spread across es, de, fr, pt, zh, ja (natural developer
|
|
|
|
|
|
usage, often code-switched with English tech nouns).
|
|
|
|
|
|
- Confusion pairs to mine deliberately in the `boundary` slice, several dozen each:
|
|
|
|
|
|
writing↔backendImpl (docs about APIs), quickFix↔refactor (renames), review↔debugging
|
|
|
|
|
|
(why-questions), planning↔backendImpl ("plan and implement"), frontendImpl↔quickFix
|
|
|
|
|
|
(small UI tweaks), review↔writing (explain vs summarize).
|
|
|
|
|
|
|
|
|
|
|
|
## Anti-patterns (rejected in QA)
|
|
|
|
|
|
|
|
|
|
|
|
- Label leakage: prompts must never contain the label word used AS a label hint
|
|
|
|
|
|
("refactor this" is fine and common; "this is a refactor task:" is not).
|
|
|
|
|
|
- Template smell: recycled sentence skeletons with one noun swapped; enumerated
|
|
|
|
|
|
"Task 47:" prefixes; uniform lengths; every example ending in a period.
|
|
|
|
|
|
- Impossible labels: prompts a human labeler couldn't defend from the text alone
|
|
|
|
|
|
(except the `vague-eval` slice, where that's the point).
|
|
|
|
|
|
- Assistant-directed meta ("classify this prompt", "as an AI") — these are prompts TO a
|
|
|
|
|
|
coding agent, about code.
|
2026-07-29 21:01:20 -07:00
|
|
|
|
- Mid-conversation replies: anything that only makes sense as turn 2+ of a thread —
|
|
|
|
|
|
reacting to an assistant's previous answer ("yes do option 2", "that didn't work, try
|
|
|
|
|
|
again", "same error as before", "looks good, ship it", "no, the OTHER function").
|
|
|
|
|
|
Referencing prior *work* in a fresh chat is fine (see the intro); referencing a prior
|
|
|
|
|
|
*turn of this conversation* is not, because the classifier never sees those.
|
2026-07-29 19:00:34 -07:00
|
|
|
|
|
|
|
|
|
|
## Process
|
|
|
|
|
|
|
|
|
|
|
|
Produce the dataset in batches of 200 lines. Before each batch, silently pick a fresh
|
|
|
|
|
|
combination of 3 domains + 2 registers + 1 confusion pair to emphasize, so no two
|
|
|
|
|
|
batches have the same texture (do not print your picks — JSONL lines only). After each
|
|
|
|
|
|
batch, self-check against the anti-patterns and the slice/label distributions, and fix
|
|
|
|
|
|
violations before emitting. Across the full run, keep primary labels within ±15% of
|
|
|
|
|
|
uniform across the 8 classes (the vague-eval slice is exempt). Target total: 8,000
|
|
|
|
|
|
lines. Do not number examples. Do not wrap output in markdown fences. JSONL only.
|