From ef1fc5af1f9cf6b493a766d3d11c3ebd594cd94a Mon Sep 17 00:00:00 2001 From: Andrew Blakeslee Moore Date: Wed, 22 Jul 2026 09:14:35 +0000 Subject: [PATCH] Unwrap AFM tool-calling scaffolding; retune two prompts for macOS 27 beta 4 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The Foundation Models seed in macOS 27 beta 4 (26A5388g) answers a plain `respond(to:)` as if it were driving a tool-calling harness, so free-form replies now arrive wrapped: a bare JSON object, a fenced ```json block, a `tool_call: {…}` literal, `[start_x]`/`[No tools needed]` markers, or a persona preamble ahead of the answer. Every consumer takes the reply's first non-empty line, so the wrapper landed verbatim in chat titles and collapsed summaries — dev even carries a commit named "Nucleic: {task:Update Website Hero Text}". Guided generation (@Generable) still binds its schema and is untouched. Add ModelOutput to NucleicProtocol — the one layer the Mac app, core and the iOS client all link — and run it at each provider's text choke point, before the existing line-oriented parsing. Plain replies pass through byte-identical; an all-scaffolding reply yields nil, which is the "no result" every caller already handles by falling back to its heuristic, so no new failure path. DelegatedIntelligence unwraps again on receipt: the mesh is mixed-version and the agent-CLI backend never unwraps. Also retune the two prompts the seed broke, measured against the live on-device model rather than by inspection: - classifyTurn: the old wording let the agentically-tuned seed reason about what the agent should do NEXT, so it answered AWAITING for plainly finished turns — 12/17 on a labeled set (3 samples, majority vote) with 4 false AWAITING, which silently stops autoship. Reframing it as "a classifier, NOT an assistant" scores 17/17 with none. Guided generation was tried here and is worse (81% struct, 68% enum). Applied to both hand-duplicated twins. - sessionKeywords: 9 of 9 runs returned a JSON/tool-call envelope, one inventing a fake Package.swift to "search". Anchoring extraction to terms already in the input returns 9 of 9 clean comma lists. Tests pin the captured beta 4 shapes verbatim and the classifier clauses that earned the accuracy, since they read like boilerplate. Co-Authored-By: Claude Opus 4.8 (1M context) --- .../Models/PhoneIntelligenceExecutor.swift | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/NucleicRemote/NucleicRemote/Models/PhoneIntelligenceExecutor.swift b/NucleicRemote/NucleicRemote/Models/PhoneIntelligenceExecutor.swift index 23ee553..4dde2a4 100644 --- a/NucleicRemote/NucleicRemote/Models/PhoneIntelligenceExecutor.swift +++ b/NucleicRemote/NucleicRemote/Models/PhoneIntelligenceExecutor.swift @@ -69,8 +69,14 @@ enum PhoneIntelligenceExecutor { let deadline = request.deadlineSeconds ?? 60 let text = await raceAgainstDeadline(seconds: deadline) { await gate.run { - try? await LanguageModelSession(model: model, instructions: job.instructions) - .respond(to: job.prompt).content + // Strip the tool-calling envelope the macOS 27 / iOS 27 model wraps free-form + // replies in before it goes on the wire — the host parses this text with the + // same line-oriented heuristics as a local generation, so it must arrive + // already unwrapped. A reply that was nothing but scaffolding yields nil, which + // answers `error` below and drops the host straight to its heuristic. + ModelOutput.usableText( + try? await LanguageModelSession(model: model, instructions: job.instructions) + .respond(to: job.prompt).content) } } if let text, !text.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty {