Skip to content

CLI reliability: correct execution, safe mutations, and trustworthy validation #554

Description

@spences10

Goal

Fix demonstrated failures in existing CLI tooling: correct execution, safe resource mutations, and trustworthy validation. Deliver small, independently reviewable fixes rather than another broad rewrite.

This tracker follows a local source audit and isolated reproductions at commit 57f1407142e6318a5d0c935127b6461a5e87526a. Implementers must recheck current source before making changes. Dependency updates are not a prerequisite for recording or prioritizing these issues.

Scope

  • CLI argument handling and output mode selection.
  • Safe skill deletion.
  • LSP diagnostic results that distinguish clean from incomplete validation.
  • Harness reporting and enforcement of changes made after its baseline.

First delivery batch

These four child issues define the initial delivery scope. Each can ship independently. Prioritize skill deletion and CLI correctness; neither requires a new workflow framework.

Delivery requirements

  • Reproduce each failure in a regression test before fixing it.
  • Test the affected entry point, not just an isolated helper when the failure occurs between components.
  • Run focused package checks/tests and source diagnostics. Run wider gates when shared wiring, manifests, or tooling change.
  • Preserve existing public behaviour outside the stated correction, along with trust, confirmation, accounting, and provenance requirements.
  • Report actual validation results and limitations. A timeout or missing result must not become a success claim.
  • Keep each fix independently reviewable. User approval remains required for commits, Changesets, and releases.

Issue map beyond the first batch

These are triage candidates, not additional implementation commitments or completion requirements for this tracker.

Finding Next treatment Related history
LSP request-ID collision; MCP unsettled requests and cancellation Separate focused corrective issues #115 (MCP lifecycle), #59 (failure-mode tests)
Settings lost updates; preset mutation bypass Separate corrective issues #288; #499, #501, #502
External skills installation suppresses resources Verify through runtime integration before filing the precise defect #289, #293
Redaction disagreement; Svelte remediation blocking Separate corrective issues #372, #517, #518; #182
Harness path-matcher disagreement and obsolete instructions Separate focused fixes; decide creation-tool/API simplification before implementation #524; #471 for upstream convergence only
Context-store closure; provider refresh after shutdown Focused lifecycle fixes, not performance projects #77
Repeated footer, skill-profile, and settings collection Measure operation counts and representative latency before selecting refactors #428 (preserve canonical accounting)
Duplicate interfaces, legacy formats, private TUI coupling, repeated CI work Later cleanup with explicit compatibility boundaries #170, #293, #294, #499
Operational observability races and error classification Separate runtime correctness follow-ups Supporting tooling, not marketing
Third-party MCP adapter integration Investigate the existing user request through existing APIs #547

Existing product-decision tracks

Do not create duplicate epics or make this reliability work depend on new capabilities:

Deferring expansion while reliability work proceeds is a recommendation, not a change to those issues or their approval state. #524 records the reason to avoid converting every candidate into a committed subsystem.

Non-goals

  • apps/web: marketing is entirely excluded.
  • New goal graphs, subagent runtimes, Team Mode schedulers, or Factory capabilities.
  • Broad package mergers, transport replacements, or unconditional removal of compatibility paths.
  • Performance refactors without measurements.
  • Reopening completed epics wholesale. Corrective issues should link their predecessors and describe the remaining failure precisely.

Completion

The first batch is complete when all four child issues meet their acceptance checks with recorded evidence. Later candidates require separate triage and approval; they must not silently expand this tracker's completion scope.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    epicLarge parent issue grouping related work

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions