Repository navigation
Replace Codex MCP dispatch with a Quest CLI runner - #177
Conversation
Use explicit cached or API-key authentication for Claude-to-Codex tasks and Quest roles, preserving native Codex delegation and installed paths. Correct installer/preflight guidance and validate artifacts and cleanup. Refs #176. Required installed live acceptance is tracked separately. Quest/Co-Authored by Co-Authored-By: GPT-6 Astra <noreply@openai.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Use a TOML table value supported by the actual CLI override parser. Preserve unrelated MCP entries and cover dashed and dotted names with real configuration parsing regressions. Refs #176 Co-Authored-By: GPT-6 Astra <noreply@openai.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Co-Authored-By: KjellKod <KjellKod@users.noreply.github.com>
There was a problem hiding this comment.
3 issues found across 34 files
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="scripts/quest_preflight.sh">
<violation number="1" location="scripts/quest_preflight.sh:318">
P2: When the codex runner probe fails (timeout, crash, or non-JSON output), the fallback discards the runner's captured stderr and emits a generic "Reinstall Quest and rerun preflight" warning with codex_auth_reason set to "not_checked". In this narrow failure mode the real diagnostic is the most valuable output, but it is never surfaced. Print the captured stderr into the warning array so the operator can act on the actual failure instead of being told to reinstall.</violation>
</file>
<file name="scripts/quest_runtime/codex_runner.py">
<violation number="1" location="scripts/quest_runtime/codex_runner.py:641">
P2: In `_validate_handoff`, the codex handoff is validated against the `next` and `artifacts` keys, but the canonical handoff contract in `.ai/schemas/handoff.schema.json` defines `next_role` (required) and `artifacts_written` with `additionalProperties: false`. If the Codex agent produces a schema-compliant handoff, every role output is rejected as `malformed_output` and the retry ladder loops forever; if it produces the `workflow.md` format, the schema still rejects the file. The new validator should read the same field names the agent is instructed to write (`next`/`artifacts` per `.skills/quest/delegation/workflow.md` line 455) or the schema's names, and the repo's two handoff contracts should be reconciled before this validation is relied on.</violation>
<violation number="2" location="scripts/quest_runtime/codex_runner.py:761">
P2: Codex context-health entries are written without `model=` and `effort=` fields, even though the documented Codex log contract for this workflow requires those fields (with legacy compatibility when present). `append_context_health_log` in claude_runner.py has no `model`/`effort` parameters, so the new Codex role path cannot emit them. Add optional `model`/`effort` params to the helper and pass the saved `model` and `effort` in this call so Codex lines carry the effective settings.</violation>
</file>
Tip: cubic can generate docs of your entire codebase and keep them up to date. Try it here.
Re-trigger cubic
Reject invalid timeout and model inputs before dispatch. Preserve known execution and authentication metadata when receipt or role logging fails. Return structured preflight diagnostics without exposing captured secrets, keep optional installer scans nonfatal, and align authentication examples and validation schema. Quest/Co-Authored by Co-Authored-By: GPT-6 Astra <noreply@openai.com> in collaboration with KjellKod
There was a problem hiding this comment.
All reported issues were addressed across 12 files (changes from recent commits).
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
Keep the runner aligned with Quest fix_iteration, which starts at zero. Cover both reviewer slots with fresh artifact, receipt and logging checks. Quest/Co-Authored by Co-Authored-By: GPT-6 Astra <noreply@openai.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> in collaboration with KjellKod
There was a problem hiding this comment.
0 issues found across 3 files (changes from recent commits).
Requires human review: Auto-approval blocked by 7 unresolved issues from previous reviews.
Re-trigger cubic
The Codex plan-reviewer example prompt allowed NEXT: null when blocked, but the handoff contract and runner require next: arbiter, so an honest blocked handoff was rejected as malformed and retried once for nothing. The installer overwrites files with mv and only restores the executable bit for EXECUTABLE_FILES. session-start.sh was never listed, so any update that rewrote the hook left it non-executable and Claude's SessionStart hook failed with permission denied. Add it to the list and assert every manifest hook under .claude/hooks is covered. Both surfaced by Codex review on a consumer install of this branch. Quest/Co-Authored by Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
There was a problem hiding this comment.
All reported issues were addressed across 3 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
|
This works, I haven't done all the validation yet, but I've been running this, and it works fine. I will do PR-shepherd. And all the validation data. This is not a high priority right now, but if people want to use this, they can. |
⭐ Why this matters
Restore Claude-to-Codex delegation through supported CLI commands using an existing ChatGPT login. The same runner accepts explicit API-key credentials, without requiring a plugin or MCP shim.
Summary
Replace removed Codex MCP dispatch with a Quest-owned Python runner around the installed Codex CLI. Claude can use cached ChatGPT login or explicitly selected API-key authentication. Codex-to-Codex Quest roles continue through native subagents. References #176.
Changes
Shepherd update: finite timeout/model validation, retained execution/auth metadata on output failures, structured nonsecret preflight errors, nonfatal targeted legacy scans, and consistent authentication docs/schema. All 18 Cubic findings triaged: 14 fixed, 3 rejected with evidence, 1 telemetry extension deferred.
Live acceptance fix: accept initial code-review iteration
0, matching Quest state. Both reviewer-slot regressions failed before the fix and pass afterward; negative iterations still fail before dispatch.Validation
Candidate: ddf9426. Codex 0.154.0, Claude 2.1.269. No merge.
Both live dispatch paths have been exercised, with explicit revision boundaries. Actual Claude Build was explicitly approved and completed, followed by dual code review, arbiter and fixer. Live testing exposed a real defect: the first review uses
fix_iteration=0, which the old runner rejected. The host's--iter 1workaround is recorded as FAIL for canonical dispatch, not accepted as proof of correctness.The guided human walkthrough is below. Existing test results and their revision boundaries are available here:
Completed validation evidence and known limitations
Evidence below is relative to
.quest/codex-cli-dispatch_2026-09-12__1152/validation/in the implementation workspace. These are retained local artifacts, not publicly attached CI artifacts. The test plan below gives reproducible commands and exact expected outcomes.Exact target check: regular nonsymlink
hello-world.txt,12bytes, hex68656c6c6f20776f726c640a, SHA256a948904f2f0f479b8f8197694b30184b0d2ed1c1cd2a1ec0fb85d299a192a447. Claude target absent before the explicit user approval and Build transition; builder write followed afterward. The independent byte checker was supervising Codex, not a human-run shell command. The generated plan's human-only wording is preserved as a procedural deviation; the original user delegated validation methods. No human execution or waiver is claimed.Other disclosed limits: external Quest symlink journal discovery fails in existing completion tooling; artifacts remain available. Some role bootstrap reads used commands outside advisory role Bash lists. No host permissions or unrelated tools were disabled. Actual findings retry paths were exercised, including same-reviewer repair with prose unchanged. Runtime settings come from session metadata, never model self-identification.
Test plan
Start here: what are we checking?
Quest coordinates AI helpers through plan, review, human approval, build, and review again. The assistant running the conversation is the lead. A Build approval is permission to create the requested file, after the plan has been reviewed.
Two reviewers check the work independently. An arbiter resolves disagreements between them. Quest saves plans, review results and execution records in a
.questfolder.This PR repairs how Claude asks Codex for help. The walkthrough checks three things:
/gptUse temporary test folders, not a working project. These checks make real model calls. The two full Quests involve several planning/review rounds; a single role may run for up to 30 minutes. A quiet terminal alone is not a failure.
1. Prepare the two test folders
Open Terminal A and run:
Expected: each command is found, and the last command reports
Logged in using ChatGPT. If it does not, runcodex login, finish the browser login, then check again. If Claude needs sign-in, launchclaude, finish sign-in, and exit back to the shell. No API key is needed for this walkthrough.Expected: the final
PASS: installed ...line contains the revision above, followed by two launch commands with real folder paths. Keep Terminal A open, because later checks use itsquest176_labvariable.If setup stops: do not continue. If the PR branch has moved, the revision guard intentionally stops the test so a different version cannot be mistaken for this one. Record the failed command and error. Do not switch to
mainto get past it.2. Ask Claude to use Codex once
Expected: answer
323, plus a successful tool call toscripts/quest_codex_runner.py task. Its receipt should sayresult_kind: complete,auth_mode: cached,auth_kind: chatgpt, andcleanup: complete.Fail: Claude only answers
323, uses a Codex MCP tool, cannot find/gpt, or reports a failed runner. The answer alone does not prove delegation. Save the error/receipt; do not install a shim or change authentication to make this test pass.3. Run a Claude-led Quest, with a real approval pause
If Quest asks which models to use, keep the defaults listed in the prompt.
Expected pause: Claude presents a reviewed plan and asks whether to Build. It may create planning records under
.quest; it must not createhello-world.txtyet.In Terminal A, select the Claude target:
quest176_target="$quest176_lab/Claude consumer/hello-world.txt"Run the Before approval check below. Only if it passes, send this in the Claude conversation:
Expected after approval: Claude has Codex create the file, then runs two code reviews. Run the After Build check below when the file has been created. If the assistant pauses for that check, paste the actual PASS output back into the conversation so review can continue.
4. Repeat with Codex leading
In Terminal A, select the Codex target:
quest176_target="$quest176_lab/Codex consumer/hello-world.txt"At Codex's approval pause, run Before approval. Then send
Build approved for this hello-world test Quest.in the Codex conversation. Run After Build when the file is created, and let both code reviews finish.Expected: the same exact file, but the Codex roles appear as native helper-agent calls. A terminal command that starts another Codex process is not a substitute for native delegation.
File checks, use these for each Quest
Expected:
PASS: target absent before Build approval. If the file already exists, stop that test and save the evidence. Deleting it afterward does not restore proof of the approval gate.Expected:
PASS: exact hello world plus one newline, 12 bytes. This checks the newline too; seeing the words in an editor is insufficient.Finish: save a result that someone else can verify
Human sign-off: both files pass, both approval pauses happened before file creation, both pairs of code reviews finished, and the report points to real dispatch evidence. An assistant saying “all good” without those records is insufficient. No need to understand every JSON field or rerun CI yourself. If evidence is unclear, leave that row NOT VERIFIED and attach the report for a maintainer to inspect.
When blocked: record the route being tested, last command/tool error, installed revision and report path. Do not paste credentials. Keep the temporary folders until the review is finished.
Advanced checks for maintainers: external installations, upgrades, API keys and automated regressions
The walkthrough above uses the ordinary inside-repository installation. The following checks cover additional layouts and authentication. They are separate from the three human checks above. API-key inference remains NOT VERIFIED unless someone with an authorized key observes a successful real call.
These commands use the
quest176_labandquest176_shavariables in Terminal A. Refer to setup above for the pinned installer download and revision guard.External installation, non-Git folders and upgrades
In Terminal A, install Quest outside a repository. This folder holds the helper; a different folder will receive the output file:
Then create a separate target repository and run the installed helper:
Expected: current completed receipt, file only in the separate target, artifacts outside both roots, exact byte check passes. Repeat against Non Git target: without --allow-non-git require nonzero exit and absent target; with that flag and a fresh output directory require successful inference and exact bytes. Verify actual runtime cwd/sandbox/writable artifact root, not only argv.
For the legacy upgrade baseline, the installer accepts a branch name, not a commit SHA. The observed baseline came from main while main pointed at 319494e. Reproduce it only while this guard passes:
If main moved, stop this legacy reproduction. Do not label a newer installation as the old baseline. Existing recorded baseline receipts remain historical evidence.
In Legacy custom, add consumer-owned.txt and an unrelated MCP entry to .opencode/opencode.json; record SHA256 hashes of both before upgrading. Keep Legacy pristine unchanged. In each directory, download the candidate installer from the pinned candidate URL in step 1, run it with --branch quest/codex-cli-dispatch --force using the isolated Upgrade home/PATH above, then require .quest-version to equal the candidate SHA.
Expected: pristine managed OpenCode Codex MCP removed; customized configuration and consumer-owned.txt retain their exact hashes, with targeted migration guidance instead of blanket deletion. Actual before/after evidence: distribution/final-legacy-preservation-before.json and distribution/final-preservation-after.json.
These integration checks do not substitute for the Claude-led and Codex-led Quests above.
API-key inference without cached login, conditional
Only run when an authorized CODEX_API_KEY or OPENAI_API_KEY is already available. Never paste/log it. In Terminal A, use the installed Claude test folder. This is a direct runner check, separate from the Claude conversation:
PASS requires: real successful inference, auth_mode=api-key, auth_kind=api-key, matching runtime model/effort and exact api-probe.txt bytes, despite no prior CLI login in that new home. A dummy-key endpoint or readiness probe is not successful OpenAI inference. No authorized key means NOT VERIFIED.
Automated checks, optional local reproduction of CI
The five schema regression cases require the AJV CLI with draft2020 support; they skip if unavailable. All five passed locally on this head.
The two tests in tests/integration/test_codex_mcp_overrides.py require a real installed Codex CLI, otherwise they skip. Local PASS on Codex 0.154.0 is recorded separately. Deterministic tests cover credential isolation, stdin/argv propagation, malformed/stale/missing outputs, process-tree timeout/cancellation, retry rules and unavailable/mismatched native controls.
Notes
No merge. Both full host workflows have run with explicit Build approvals. The failed initial-review path was fixed, then passed a separate actual Claude-launched installed replay. Retained older live proof and later source/installed equivalence are explicitly separated. Live API-key inference remains conditional and unverified.
Codex acceptance planning used f649beb; the unchanged reviewed plan was built and code-reviewed on then-current candidate 0ec404d. Executable dispatch arguments and tool names were inspectable throughout; some collaboration prompt/message bodies were encrypted. Builder and native reviewer bootstrap read-only commands exceeded their declared role lists. External-symlink journal discovery failed; a manual canonical journal and installed --skip-journal archive completed the workflow. These limits are recorded, without claiming strict role allowlist compliance or global process cleanup.
Normal host sessions added Codex project registrations; structured comparison confirms MCP settings and model/effort unchanged. Consumer upgrade sentinels remain byte-identical. Existing reverse-direction Claude probe respawn behavior is recorded separately, without expanding this migration.
▐▛███▜▌ ▝▜█████▛▘ ▘▘ ▝▝ Quest/Co-Authored by Co-Authored-By: GPT-6 Astra <noreply@openai.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> in collaboration with KjellKod