# Synthetic-data generation prompt for the purpose classifier The prompt below is fed verbatim to a frontier-model agent to produce training/eval data per docs/PURPOSE_CLASSIFIER.md §4.1. Record the generating model, date, and batch topics in the generation manifest alongside the output. The 92 shipped fixtures (Tests/NucleicCoreTests/Fixtures/purpose-prompts.json) are eval-only and must NOT be pasted into the generator's context (contamination). --- You are generating a labeled dataset of prompts that software developers type into a coding-agent app (like Claude Code) to START a new chat. Each example is the OPENING message of a fresh chat — never a reply inside an ongoing conversation — labeled with what the prompt is FOR. This matters: the classifier runs exactly once, on the first message, to pick the chat's model (which then stays fixed for the chat's whole life), so the training distribution must be first-messages only. A first message may still reference prior work the way developers really do ("continuing from yesterday's auth refactor, …", "picking up the payment-flow bug again"), because people routinely start fresh chats mid-project. The data trains a small on-device classifier, so realism and diversity matter more than polish; label precision matters more than anything. ## Labels (choose the primary purpose; definitions are exhaustive) - `planning` — asking for architecture, design docs, RFCs, migration strategy, roadmaps, breaking work into milestones. The deliverable is a PLAN or DESIGN, not code. - `backendImpl` — implementing server/API/data/algorithm/system/CLI code: endpoints, schemas, migrations, queues, caches, auth flows, parsers, background jobs. - `frontendImpl` — implementing UI: views, components, styling, layout, animation, themes, screens, visual polish. If the deliverable is something you SEE, it's frontend. - `quickFix` — a typo, version bump, config tweak, flag flip, one-liner, or a small contained bugfix the author already understands. Small scope, known change. - `refactor` — restructuring without behavior change: rename/extract/split/consolidate/ dedupe/decouple/simplify. The author expects identical behavior after. - `debugging` — diagnosing a failure the author does NOT yet understand: crashes, stack traces, regressions, flaky tests, hangs, leaks, wrong output, "why does X happen". - `review` — reading/judging/explaining EXISTING code or designs: code review, audits, "what does X do", "is this safe", comparisons, walkthroughs. No code changes requested. - `writing` — producing prose: docs, READMEs, commit messages, PR descriptions, release notes, changelogs, summaries, translations, doc comments. Boundary rules (apply in this order when two labels tempt you): 1. "Fix" + author already knows the change → `quickFix`. "Fix" + cause unknown / symptoms described → `debugging`. 2. Rename/restructure "across the codebase" or preserving behavior → `refactor`, even though a single rename in one file reads as `quickFix`. 3. Docs/comments/prose about code → `writing`, even when the subject is an API or backend concept ("update the API docs" is `writing`). 4. "Plan/design/architect X" → `planning` even when X is backend or frontend work. "Plan and implement X" → primary is `planning`, secondary is the implementation label. 5. A pure question about existing behavior → `review` unless something is BROKEN, then `debugging`. ## Output format — strict JSONL, one object per line, no commentary {"prompt": "...", "purpose": "", "secondary": "