Skip to content

Commit d3f3c57

Browse files
feat: recover empty and tool-end finalize (v0.7.3)
Detect synthetic blocked JSON and empty finalize, run finish-steer recovery passes, and fall back to provisional done when disk/tool evidence already proves writes landed. Add Python/Go knowledge bars for language expectations.
1 parent 73cb78e commit d3f3c57

13 files changed

Lines changed: 346 additions & 173 deletions

File tree

‎Formula/slmcode.rb‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -9,7 +9,7 @@
99
class Slmcode < Formula
1010
desc "Coding harness for SLMs and any OpenAI-compatible LLM"
1111
homepage "https://unicolab.ai"
12-
version "0.7.2"
12+
version "0.7.3"
1313
license "MIT"
1414

1515
on_macos do

‎Makefile‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
MODULE := github.com/UnicoLab/slmcode
22
BIN := slmcode
3-
VERSION ?= 0.7.2
3+
VERSION ?= 0.7.3
44
PREFIX ?= $(HOME)/.local
55
GIT_COMMIT := $(shell git rev-parse --short HEAD 2>/dev/null || echo unknown)
66
BUILD_TIME := $(shell date -u +%Y-%m-%dT%H:%M:%SZ)

‎TODO.md‎

Lines changed: 0 additions & 108 deletions
Original file line numberDiff line numberDiff line change
@@ -1,109 +1 @@
1-
## Fix agents in the studio
2-
**Done.** `GET /api/agents/{id}` returns full config including built-in `system_prompt`. Studio loads detail on row click / Edit. TUI `/agent show|edit` seeds from the same path. Verified live via Studio API.
31

4-
## Tester agent
5-
**Done.** Tester prompt/skill require real `ws_shell` execution. Soft-pass without `commands[]` is rejected. Per-task tester finish schema fixed. Finalize mandates commands. Live oMLX run executed `python main.py --help` / smoke checks.
6-
7-
## Project plan
8-
**Done.** `ParsePlanJSON` accepts string or object steps; never dumps raw JSON into Summary. Unparsed output goes to a fenced Raw appendix. Live PLAN.md stayed structured.
9-
10-
## Spec / functional code quality
11-
**Done.** Clarifier + greenfield harness (`requirements.txt` / pytest smoke) + auto tester task.
12-
Deterministic `post_worker_smoke` (`py_compile` / `go test`) before review; pre-test smoke before LLM tester;
13-
QA gate uses compileall/pytest (never `--help`). Soft-pass needs real Observation/smoke evidence.
14-
Live oMLX greenfield e2e PASS (~98s).
15-
16-
## TUI improvements
17-
**Done.** Live throttled redraw during runs + mid-run board refresh. `KindDebug` for runner internals (hidden by default). Studio Debug toggle. Compact mode default on. SSE prefers latest event under backpressure.
18-
19-
## Check the current failure and improve
20-
**Done.** Full `go test ./...` green. Offline quality e2e + live `TestLiveQualityGreenfieldPython` (oMLX) PASS — produced working `main.py` with real execution evidence.
21-
22-
## Planning / scoping (PRD interview)
23-
**Done.** Claude Code AskUserQuestion + pi-clarify style interviewer (`clarify_mode`: auto|ask|off),
24-
locked PRD in CONTEXT, Studio ask modal + `POST /api/clarify/answer`, post-split `scope_judge`
25-
enriching every task with concrete acceptance/checklist before execute.
26-
27-
## Claude Code gap ports
28-
**Done.** Plan approve gate · real CONTEXT `/compact context` · interactive shell ask ·
29-
wave rewind snapshots · hooks.json Pre/PostToolUse · thin read-only MCP (`mcp_call`) ·
30-
wired `auto_approve`.
31-
32-
## little-coder SLM harness ports
33-
**Done.** Write-guard · read-before-edit · shell-write guard · tool/knowledge inject ·
34-
quality-monitor · reserved device names · path normalize · tool-output truncate.
35-
36-
## Code quality > giant LLMs
37-
**Done.** Numbered ws_read + offset/limit + auto-trim · fuzzy edit recovery · static gate ·
38-
require smoke · finalize-warn + thinking-budget · text tool-call detection · JS smoke ·
39-
knowledge cards · stricter reviewer/corrector · cooler temps.
40-
41-
**Also done.** Over-edit guard (refuse whole-file-style edits) · no-op edit refuse ·
42-
claims gate (hallucinated `files_changed`) · worker self-critique on weak output ·
43-
algorithm cheat-sheets (binary search, DP, two pointers, BFS/DFS, backtracking, sort).
44-
Config: `claims_gate`, `worker_critique`, `over_edit_guard` (default ON).
45-
46-
## little-coder quality ports (round 4)
47-
**Done.** Worker multipass when `think_passes≥2` (incomplete JSON → critique/refine) ·
48-
finalize mid-run steer on ReAct resume · react context-watchdog (conversation compact at
49-
80% of `max_context_kb` with #68 hysteresis). Config: `react_compact`,
50-
`react_compact_at_percent` (default ON / 80).
51-
52-
## little-coder quality ports (round 5)
53-
**Done.** Mid-ReAct loopguard (refuse verbatim repeated tool calls via wrappers) ·
54-
output-parser extract + arg-rich text-tool nudge · edit/write/patch failure recovery
55-
cards in tool results. `quality_monitor` now also wraps coding tools at register time.
56-
57-
## little-coder quality ports (round 6)
58-
**Done.** Remaining knowledge cards (hash/tree/BFS-state/io-wrapper/string-rules/zipper) ·
59-
bash SAFE_PREFIXES whitelist (`shell_whitelist`, `shell_allow`) · thinking-budget hard
60-
abort recovery (post-turn) · first-write-wins file checkpoints · per-model profiles
61-
(`model_profiles`) · malformed_args quality detection.
62-
63-
## UX + eval (v0.6.0)
64-
**Done.** `KindIntervention` / `KindTurn` SSE · TUI banners + turn meter + progress strip ·
65-
prompt history · `/clear` · `/plan` · richer `/help` · Studio intervention banner + turn chip ·
66-
read trim by context % · `slmcode eval` + `pkg/eval` harness · Studio UI markers aligned.
67-
68-
## Quality: no more false success on placeholder scaffolds (TestSLMs)
69-
**Done.** Root cause: `promoteBoardOnQAGreen` + weak `compileall` QA rubber-stamped
70-
escalated/blocked tasks; static gate missed `# Placeholder implementation`; existence-only
71-
acceptance / already-satisfied skipped real work; Studio had no escalate intervention.
72-
73-
Fixes:
74-
- Never promote escalated / `status=blocked` / missing files / static stubs on QA green
75-
- Weak QA (`compileall` / `py_compile`) does not clear tester rejection or escalate backlog
76-
- Run success requires no escalated tasks left
77-
- Static quality catches Placeholder comments + stub constant returns
78-
- Review-time static insurance + Studio `KindIntervention` on escalate / stub refuse
79-
- Vague acceptance rejects "exists and contain…", "tool evidence", `pytest --collect-only`
80-
- Architect/splitter/reviewer/tester/worker prompts require functional code, not stubs
81-
- Greenfield harness now catches "setup/template/folder structure" + LangGraph class agents
82-
(auto-adds requirements.txt, main.py, pytest) + LangGraph knowledge card
83-
- Studio Live: overview grid (status / activity / enrich / deps) + dedicated `LiveLogs` panel
84-
- Studio pipeline header redesigned: Prepare/Design/Build/Verify/Finish groups, track fill,
85-
live agent/task, board %, harness intervention chip
86-
- Static gate rejects `from langgraph import Graph`; clarifier/PRD defaults force runnable
87-
LangGraph acceptance; e2e `TestLangGraphTemplateQueryPath` locks the harness path
88-
- Placeholder specialist (`@placeholder`) + project-wide stub scan after execute (polish phase);
89-
fills real code or flags precise `path:line` gaps; Studio continue-ask when retries/QA
90-
exhausted (`continue` / `stop` / `flag_only`) via `/api/continue/*`
91-
92-
## Pipeline config + reference-quality bar (v0.7.0)
93-
**Done.** Config-driven pipeline (`.slmcode/pipeline.yaml`) with phase agent bindings,
94-
execute-loop reviewer/corrector, and insertable slots (before/after/replace) with full
95-
prompt/params control. Studio Pipeline tab + dynamic progress header follow the same config.
96-
`GET/PUT /api/pipeline`, custom agents usable in specialist mode and board roles.
97-
98-
Also: project completeness reference bar (LangGraph/FastAPI/CLI), real-query eval suite,
99-
loop SSE (`kind=loop`) with continue/abort UX, preserve harness tasks when capping to 8,
100-
docs for pipeline + updated agents/studio/config/architecture.
101-
102-
## Runnable quality bar + escalate HITL (v0.7.1 / v0.7.2)
103-
**Done.** Push SLM output past "compiles": greenfield Python QA → pytest; whitelisted
104-
acceptance commands run after workers; critique until smoke/static/acceptance green;
105-
syntax-only QA cannot alone mark run success.
106-
107-
Mid-execute escalate pauses with Studio modal + `GET/POST /api/escalate/*`;
108-
TUI `/escalate re_scope|retry|mark_done|abort`. On timeout **@escalate** (or
109-
`escalate_timeout_agent`) decides — not blind re_scope.

‎cmd/slmcode/version.go‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@ package main
44
//
55
// go build -ldflags "-X main.Version=0.5.0 -X main.SourceRoot=/path -X main.GitCommit=abc -X main.BuildTime=…"
66
var (
7-
Version = "0.7.2"
7+
Version = "0.7.3"
88
SourceRoot = "" // absolute path to the slmcode checkout used to build this binary
99
GitCommit = "unknown"
1010
BuildTime = "unknown"

‎docs/changelog.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,15 @@
11
# Changelog
22

3+
## v0.7.3 — Incomplete finalize recovery
4+
5+
### Highlights
6+
- Detect empty finalize and synthetic `model ended on a tool call` blocked JSON
7+
- Up to two finish-steer corrector passes (demand status JSON, stop tool chains)
8+
- Provisional done from disk/tool evidence when finalize still fails
9+
- Knowledge cards: Python / Go project bars for language expectations
10+
11+
---
12+
313
## v0.7.2 — Escalate timeout → SLM arbitrator
414

515
### Highlights

‎docs/install.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -75,7 +75,7 @@ One-liners fetch a shiny GitHub Release. Keep the compiler for contributing (or
7575
7676
```bash
7777
curl -fsSL https://raw.githubusercontent.com/UnicoLab/smlcode/main/scripts/install-remote.sh \
78-
| bash -s -- --version v0.7.2
78+
| bash -s -- --version v0.7.3
7979
```
8080
8181
=== "🪟 PowerShell"

‎pkg/augment/augment.go‎

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -180,6 +180,34 @@ After writes: install deps if needed, run smoke, then status=done.`,
180180
- ZERO stub markers or fake {"output":"run_result"} returns
181181
After writes: pip install -r requirements.txt && python -m pytest -q && python main.py`,
182182
},
183+
{
184+
Topic: "Python Project Bar", TokenCost: 120,
185+
Keywords: []string{
186+
"python", "pytest", "pip", "requirements.txt", "pyproject", "venv",
187+
"fastapi", "django", "flask", ".py",
188+
},
189+
RequiresTools: []string{"ws_write", "ws_shell"},
190+
Body: `Python expectations (unless the repo already disagrees):
191+
- Real modules, not Placeholder/TODO stubs; prefer typed public functions
192+
- tests/ with pytest (or test_*.py next to modules); no fake green asserts
193+
- requirements.txt or pyproject.toml with pinned direct deps
194+
- Entry: main.py or package __main__; runnable smoke after install
195+
Verify: python -m pip install -r requirements.txt && python -m pytest -q`,
196+
},
197+
{
198+
Topic: "Go Project Bar", TokenCost: 120,
199+
Keywords: []string{
200+
"golang", "go mod", "go test", "go.mod", "goroutine", "pkg/",
201+
"cmd/", ".go",
202+
},
203+
RequiresTools: []string{"ws_write", "ws_shell"},
204+
Body: `Go expectations (unless the repo already disagrees):
205+
- Valid module path in go.mod; packages compile with go test ./... -short
206+
- Exported APIs documented lightly; no panic("TODO") / unimplemented stubs
207+
- Prefer table-driven tests; keep edits inside listed packages
208+
- cmd/ for entrypoints, pkg/ or internal/ for libraries when scaffolding
209+
Verify: go test ./... -short (or targeted package) before status=done`,
210+
},
183211
}
184212
return append(base, AlgorithmKnowledge()...)
185213
}

‎pkg/augment/augment_test.go‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -84,3 +84,18 @@ func TestSelectKnowledgeLangGraphClassAgent(t *testing.T) {
8484
t.Fatalf("expected LangGraph Class Agent card, got %#v", ks)
8585
}
8686
}
87+
88+
func TestSelectKnowledgePythonAndGoBars(t *testing.T) {
89+
py := SelectKnowledge("add pytest coverage for the fastapi python service", DefaultKnowledge(), 400)
90+
goK := SelectKnowledge("fix go test failures in pkg/loop", DefaultKnowledge(), 400)
91+
assertTopic := func(ks []KnowledgeEntry, topic string) {
92+
for _, k := range ks {
93+
if k.Topic == topic {
94+
return
95+
}
96+
}
97+
t.Fatalf("expected %s card, got %#v", topic, ks)
98+
}
99+
assertTopic(py, "Python Project Bar")
100+
assertTopic(goK, "Go Project Bar")
101+
}

‎pkg/loop/runner.go‎

Lines changed: 111 additions & 59 deletions
Original file line numberDiff line numberDiff line change
@@ -344,62 +344,9 @@ func (r *Runner) runWave(ctx context.Context, board *plan.Board, wave []plan.Tas
344344
if res.Iteration > 0 {
345345
r.fireTurn(t.ID, res.Iteration, roleMaxIter(role))
346346
}
347-
// SLMs sometimes emit a bare tool call / empty finalize — nudge one corrective pass.
348-
needNudge := looksLikeToolJunk(t.Output)
349-
nudgeIssue := "Finish the task and return status JSON. Do not end on a tool call. Prefer ws_edit/ws_patch (ws_read first); ws_write only for NEW files."
350-
if r.QualityMonitor && !needNudge {
351-
assess := quality.AssessResponse(t.Output, nil, nil, nil)
352-
if !assess.OK {
353-
needNudge = true
354-
nudgeIssue = quality.CorrectionMessage(assess.Reason)
355-
if strings.HasPrefix(assess.Reason, "text_tool_calls:") {
356-
if calls := quality.ParseTextToolCalls(t.Output); len(calls) > 0 {
357-
nudgeIssue = quality.TextToolNudge(calls)
358-
}
359-
}
360-
r.Log("%s quality-monitor: %s", t.ID, quality.PhraseForUser(assess.Reason))
361-
r.fireIntervention(t.ID, assess.Reason, quality.PhraseForUser(assess.Reason), assess.Reason)
362-
}
363-
}
364-
if !needNudge {
365-
if calls := quality.ParseTextToolCalls(t.Output); len(calls) > 0 {
366-
needNudge = true
367-
nudgeIssue = quality.TextToolNudge(calls)
368-
if r.AutoTextTools {
369-
nudgeIssue = "AUTO_TEXT_TOOLS: re-issue these as NATIVE tool calls immediately, then status JSON.\n" + nudgeIssue
370-
}
371-
r.Log("%s text-tool-parser: recovered %d call(s)", t.ID, len(calls))
372-
r.fireIntervention(t.ID, "text_tool_calls", "text tool calls recovered", nudgeIssue)
373-
}
374-
}
375-
if !needNudge && r.ThinkingBudget &&
376-
quality.ThinkingBudgetExceeded(t.Output, r.ThinkingBudgetTokens) {
377-
needNudge = true
378-
nudgeIssue = quality.ThinkingBudgetBreachMessage()
379-
r.Log("%s thinking-budget: exceeded — forcing commit pass", t.ID)
380-
r.fireIntervention(t.ID, "thinking_budget_exceeded", "thinking budget exceeded", nudgeIssue)
381-
}
382-
if needNudge {
383-
r.Log("%s produced incomplete finalize; running corrector once", t.ID)
384-
r.fire("agent_start", r.correctorID(), t.ID, "fix incomplete finalize", strings.Join(t.Files, ", "), "")
385-
corrIn := formatCorrectPrompt(t, plan.ReviewResult{
386-
Approved: false,
387-
Issues: []string{nudgeIssue},
388-
Summary: "incomplete finalize",
389-
})
390-
corr, _ := r.Executor.ExecuteSubAgents(ctx, []ggagent.SubAgentRequest{{
391-
AgentID: r.correctorID(),
392-
Input: corrIn,
393-
Timeout: r.Timeout, ShareState: true,
394-
}}, r.Shared)
395-
if len(corr) > 0 {
396-
r.noteUsage(corr[0], corrIn, outputString(corr[0]))
397-
if out := outputString(corr[0]); strings.TrimSpace(out) != "" {
398-
t.Output = out
399-
}
400-
}
401-
r.fire("agent_end", r.correctorID(), t.ID, "corrector finished", "", truncate(t.Output, 800))
402-
}
347+
// SLMs often empty-finalize or end on tool-junk / synthetic blocked JSON.
348+
// Recover up to 2 finish-steer passes; then provisional-done if disk proves writes.
349+
r.recoverIncompleteFinalize(ctx, &t, snapshots[i])
403350
mergeFilesChanged(&t)
404351
// Attach disk evidence hint for reviewer.
405352
if hint := r.diskEvidenceHint(t, snapshots[i]); hint != "" {
@@ -1166,9 +1113,114 @@ func outputString(res ggagent.SubAgentResult) string {
11661113
}
11671114

11681115
func looksLikeToolJunk(raw string) bool {
1169-
lower := strings.ToLower(raw)
1170-
return strings.Contains(lower, "<function") || strings.Contains(lower, "<tool_call>") ||
1171-
strings.Contains(lower, "</tool_call>")
1116+
return quality.LooksLikeToolJunk(raw)
1117+
}
1118+
1119+
// recoverIncompleteFinalize nudges the corrector (up to 2 passes) when the
1120+
// worker ended empty / on tool-junk / with GoLangGraph's synthetic blocked JSON.
1121+
// If recovery still fails but disk/tool evidence shows writes, synthesize a
1122+
// provisional done JSON so review/smoke can decide on real evidence.
1123+
func (r *Runner) recoverIncompleteFinalize(ctx context.Context, t *plan.Task, baseline map[string]string) {
1124+
if r == nil || t == nil {
1125+
return
1126+
}
1127+
const maxPasses = 2
1128+
for pass := 0; pass < maxPasses; pass++ {
1129+
reason, nudgeIssue, ok := r.incompleteFinalizeNudge(*t)
1130+
if !ok {
1131+
return
1132+
}
1133+
r.Log("%s incomplete finalize (%s); finish-steer pass %d/%d", t.ID, reason, pass+1, maxPasses)
1134+
r.fireIntervention(t.ID, reason, quality.PhraseForUser(reason), nudgeIssue)
1135+
r.fire("agent_start", r.correctorID(), t.ID, "fix incomplete finalize", strings.Join(t.Files, ", "), "")
1136+
corrIn := formatCorrectPrompt(*t, plan.ReviewResult{
1137+
Approved: false,
1138+
Issues: []string{nudgeIssue},
1139+
Summary: "incomplete finalize",
1140+
})
1141+
corr, _ := r.Executor.ExecuteSubAgents(ctx, []ggagent.SubAgentRequest{{
1142+
AgentID: r.correctorID(),
1143+
Input: corrIn,
1144+
Timeout: r.Timeout, ShareState: true,
1145+
}}, r.Shared)
1146+
if len(corr) > 0 {
1147+
r.noteUsage(corr[0], corrIn, outputString(corr[0]))
1148+
if out := outputString(corr[0]); strings.TrimSpace(out) != "" {
1149+
t.Output = out
1150+
}
1151+
}
1152+
r.fire("agent_end", r.correctorID(), t.ID, "corrector finished", "", truncate(t.Output, 800))
1153+
}
1154+
reason, _, stillBad := r.incompleteFinalizeNudge(*t)
1155+
if !stillBad {
1156+
return
1157+
}
1158+
hasWrite := r.hasRealWriteEvidence(*t, baseline) || hasToolWriteEvidence(t.Output) || hasDiskEvidenceSection(t.Output)
1159+
if !hasWrite {
1160+
r.Log("%s incomplete finalize persists after recovery (%s)", t.ID, reason)
1161+
return
1162+
}
1163+
files := append([]string{}, t.Files...)
1164+
files = append(files, parseFilesChanged(t.Output)...)
1165+
provisional := quality.ProvisionalDoneFromEvidence(uniqStrings(files), reason)
1166+
// Keep prior observations for smoke/static appendices that may already be present.
1167+
if rest := strings.TrimSpace(t.Output); rest != "" && !strings.HasPrefix(rest, "{") {
1168+
t.Output = provisional + "\n\n" + rest
1169+
} else {
1170+
t.Output = provisional
1171+
}
1172+
r.Log("%s provisional done from write evidence after incomplete finalize (%s)", t.ID, reason)
1173+
r.fireIntervention(t.ID, "provisional_finalize", "recovered finalize from disk/tool evidence", reason)
1174+
}
1175+
1176+
// incompleteFinalizeNudge returns whether the task output needs a finish-steer
1177+
// pass, plus the reason and corrective issue text.
1178+
func (r *Runner) incompleteFinalizeNudge(t plan.Task) (reason, nudgeIssue string, need bool) {
1179+
core := stripPostSections(t.Output)
1180+
if reason = quality.IncompleteFinalizeReason(core); reason != "" {
1181+
hasWrite := hasToolWriteEvidence(t.Output) || hasDiskEvidenceSection(t.Output)
1182+
return reason, quality.FinishSteerMessage(reason, hasWrite), true
1183+
}
1184+
if looksLikeToolJunk(core) {
1185+
return "ended_on_tool_call", quality.FinishSteerMessage("ended_on_tool_call", false), true
1186+
}
1187+
if r.QualityMonitor {
1188+
assess := quality.AssessResponse(core, nil, nil, nil)
1189+
if !assess.OK {
1190+
nudgeIssue = quality.CorrectionMessage(assess.Reason)
1191+
if strings.HasPrefix(assess.Reason, "text_tool_calls:") {
1192+
if calls := quality.ParseTextToolCalls(core); len(calls) > 0 {
1193+
nudgeIssue = quality.TextToolNudge(calls)
1194+
}
1195+
}
1196+
return assess.Reason, nudgeIssue, true
1197+
}
1198+
}
1199+
if calls := quality.ParseTextToolCalls(core); len(calls) > 0 {
1200+
nudgeIssue = quality.TextToolNudge(calls)
1201+
if r.AutoTextTools {
1202+
nudgeIssue = "AUTO_TEXT_TOOLS: re-issue these as NATIVE tool calls immediately, then status JSON.\n" + nudgeIssue
1203+
}
1204+
return "text_tool_calls", nudgeIssue, true
1205+
}
1206+
if r.ThinkingBudget && quality.ThinkingBudgetExceeded(core, r.ThinkingBudgetTokens) {
1207+
return "thinking_budget_exceeded", quality.ThinkingBudgetBreachMessage(), true
1208+
}
1209+
return "", "", false
1210+
}
1211+
1212+
func uniqStrings(in []string) []string {
1213+
seen := map[string]bool{}
1214+
var out []string
1215+
for _, s := range in {
1216+
s = strings.TrimSpace(s)
1217+
if s == "" || seen[s] {
1218+
continue
1219+
}
1220+
seen[s] = true
1221+
out = append(out, s)
1222+
}
1223+
return out
11721224
}
11731225

11741226
// stripPostSections removes harness-appended evidence/gate sections so JSON

0 commit comments

Comments
 (0)