You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Evaluate the extension on representative repository work before enabling it by default or expanding its orchestration scope.
Cases
A small deterministic task with one node.
Parallel-ready nodes with a shared downstream dependency.
Invalid and cyclic graph proposals.
Missing or false completion evidence.
A blocked task and deadlock recovery.
Pause, resume, branch, abort, and session restart.
No-progress and budget exhaustion under auto-continuation.
A harness-backed ready task without duplicate validation state.
Acceptance criteria
Runs complete without manual state repair or duplicate continuation.
Invalid graphs never partially persist.
Completion audits reject unsupported claims.
Operator effort and outcome quality are compared with plain Pi plus a prompt.
Cost and continuation behavior are recorded.
End with an explicit decision: keep opt-in, enable by default, revise, or retire.
Worker spawning, Team Mode dispatch, and multi-goal support remain separate decisions.
Tool-activation comparison
For representative task-completion cases, compare the existing plain-Pi baseline with the goal feature under fully active and on-demand tool exposure. Reuse the planned cases rather than adding another benchmark framework; see #471.
Match model/provider/thinking, task inputs, dependencies, working backends, resource configuration and budgets.
Record actual active tools, schema/prompt size, activation events, input/output/cache tokens, cost, operator actions and stop reasons.
Score verified artifacts separately from normal completion, budget exhaustion and process settlement; include failures and discovery/continuation overhead.
Use representative tracked-repository work, document cache effects, and repeat comparable runs before changing defaults.
Parent
Depends on
Scope
Evaluate the extension on representative repository work before enabling it by default or expanding its orchestration scope.
Cases
Acceptance criteria
Tool-activation comparison
For representative task-completion cases, compare the existing plain-Pi baseline with the goal feature under fully active and on-demand tool exposure. Reuse the planned cases rather than adding another benchmark framework; see #471.
Existing decision gates and orchestration exclusions remain unchanged.