Skip to content

Replace Codex MCP dispatch with a Quest CLI runner - #177

Merged
KjellKod merged 7 commits into
mainfrom
quest/codex-cli-dispatch
Sep 22, 2026
Merged

KjellKod merged 7 commits into
mainfrom
quest/codex-cli-dispatch

Conversation

@KjellKod

@KjellKod KjellKod commented Sep 12, 2026 •

Copy link
Copy Markdown
Owner

⭐ Why this matters

Restore Claude-to-Codex delegation through supported CLI commands using an existing ChatGPT login. The same runner accepts explicit API-key credentials, without requiring a plugin or MCP shim.

Summary

Replace removed Codex MCP dispatch with a Quest-owned Python runner around the installed Codex CLI. Claude can use cached ChatGPT login or explicitly selected API-key authentication. Codex-to-Codex Quest roles continue through native subagents. References #176.

Changes

Shepherd update: finite timeout/model validation, retained execution/auth metadata on output failures, structured nonsecret preflight errors, nonfatal targeted legacy scans, and consistent authentication docs/schema. All 18 Cubic findings triaged: 14 fixed, 3 rejected with evidence, 1 telemetry extension deferred.

  • Share one runner between Claude /gpt and Claude-led Quest roles. Use stdin prompts, argument arrays and saved model/effort/authentication settings.
  • Require valid current-attempt events, artifacts and handoffs. Preserve planner history, reviewer repair inputs and process-tree cleanup before retry. Normal role deadline: 1,800 seconds.
  • Separate installation, target and artifact roots, including spaces and intentional non-Git execution.
  • Update installer/preflight, manifest, host skills and active docs. Remove obsolete managed MCP defaults while preserving custom/unrelated configuration and existing reverse-direction dispatch.

Live acceptance fix: accept initial code-review iteration 0, matching Quest state. Both reviewer-slot regressions failed before the fix and pass afterward; negative iterations still fail before dispatch.

Validation

Candidate: ddf9426. Codex 0.154.0, Claude 2.1.269. No merge.

Both live dispatch paths have been exercised, with explicit revision boundaries. Actual Claude Build was explicitly approved and completed, followed by dual code review, arbiter and fixer. Live testing exposed a real defect: the first review uses fix_iteration=0, which the old runner rejected. The host's --iter 1 workaround is recorded as FAIL for canonical dispatch, not accepted as proof of correctness.

The guided human walkthrough is below. Existing test results and their revision boundaries are available here:

Completed validation evidence and known limitations

Evidence below is relative to .quest/codex-cli-dispatch_2026-09-12__1152/validation/ in the implementation workspace. These are retained local artifacts, not publicly attached CI artifacts. The test plan below gives reproducible commands and exact expected outcomes.

Check Status Observed result and revision Evidence
Automated checks PASS ddf9426: 1,315 Python tests; Black, manifest, handoff, model-policy and diff checks. Prior unchanged preflight/runtime/installer/orchestration suites passed. fix-3/checks.json, fix-3/pytest.txt, checks/
Focused dual review PASS Claude Opus 5 and Astra medium accepted the iteration fix with no must-change findings. fix-3/review-a.json, fix-3/review-b.json
Iteration regression PASS Both first-review slots failed before, passed after. Real subprocesses produce fresh validated artifacts, iter-0 receipts/logs; stale outputs and negative iterations remain rejected. fix-3/
Actual installed candidate PASS Installer fetched ddf9426 into normal inherited Claude fixture. Installed SHA and changed source bytes match; target, approved plan, saved settings and completed state unchanged. fix-3/installed-upgrade.json
Earlier distribution matrix PASS Actual fresh Git, legacy upgrades with preserved custom/unrelated files, separate install/target/artifact roots, spaced paths and intentional non-Git runs passed on 0ec404d. Fresh installation/import completeness and live non-Git reran on 7564e4f. Exact revision boundaries retained. distribution/, pr-shepherd-1/actual-install/, pr-shepherd-1/live-non-git/
Actual Claude /gpt PASS 7564e4f: real Skill(gpt) host c6e85c1c launched installed runner, Codex child01a0991b produced exact12bytes. Cached ChatGPT, actual Astra medium, full executable trace. This is a 7564e4f live run, not a new ddf run. claude-current-gpt-verified.json, claude-current-gpt-host-trace.json, claude-current-gpt-codex-trace.json
Claude-led full Quest plus corrected-path replay PASS Real approved Build, dual reviews, arbiter/fixer and completed state on 7564e4f; exact12bytes. Original iter0 FAIL preserved. Installed ddf replay: actual Claude host236829a5 invoked runner --iter0, child01a0992e returned fresh validated review artifacts, Astra medium/cached ChatGPT. Replay uses explicitly synthetic initial-review state copied from the completed Quest, not a new full Quest. claude-current-full-audit.json, claude-7564e4f-completion.json, fix-3/live-replay-audit.json
Codex-led full Quest PASS Real explicit approval, absent target before Build, exact12bytes after,11valid handoffs,6native Astra medium children and5reverse Claude Opus5 roles. Complete relevant executable trace: no nested Codex CLI or Codex MCP role dispatch. Planning f649beb, Build/reviews0ec404d. Final ddf source/installed native paths verified unchanged:15 key files identical to0ec404d,22 installed files checked overall. No claim of another full live Quest. codex-acceptance-completion.json, codex-quest-final-dispatch-audit.json, codex-current-head-coverage.json
Live API-key inference without cached login NOT VERIFIED Credential routing/no-login boundary tests pass. No authorized key environment available for real inference. Dummy keys do not prove OpenAI inference. matrix.json L4

Exact target check: regular nonsymlink hello-world.txt,12bytes, hex 68656c6c6f20776f726c640a, SHA256 a948904f2f0f479b8f8197694b30184b0d2ed1c1cd2a1ec0fb85d299a192a447. Claude target absent before the explicit user approval and Build transition; builder write followed afterward. The independent byte checker was supervising Codex, not a human-run shell command. The generated plan's human-only wording is preserved as a procedural deviation; the original user delegated validation methods. No human execution or waiver is claimed.

Other disclosed limits: external Quest symlink journal discovery fails in existing completion tooling; artifacts remain available. Some role bootstrap reads used commands outside advisory role Bash lists. No host permissions or unrelated tools were disabled. Actual findings retry paths were exercised, including same-reviewer repair with prose unchanged. Runtime settings come from session metadata, never model self-identification.

Test plan

Start here: what are we checking?

Quest coordinates AI helpers through plan, review, human approval, build, and review again. The assistant running the conversation is the lead. A Build approval is permission to create the requested file, after the plan has been reviewed.

Two reviewers check the work independently. An arbiter resolves disagreements between them. Quest saves plans, review results and execution records in a .quest folder.

This PR repairs how Claude asks Codex for help. The walkthrough checks three things:

Check What it proves
Claude /gpt Claude can send a small task to Codex using the installed helper.
Claude-led Quest Claude can use Codex to create and review a file, after approval.
Codex-led Quest Codex can delegate to its built-in helper agents, without calling itself through the old MCP connection.

Use temporary test folders, not a working project. These checks make real model calls. The two full Quests involve several planning/review rounds; a single role may run for up to 30 minutes. A quiet terminal alone is not a failure.

1. Prepare the two test folders

  • Prerequisites: macOS or Linux, Git, Python 3.10+, curl, jq, and working Claude and Codex CLIs. Both accounts need access to the configured models, Claude Opus 5 and GPT-6 Astra. Use the normal installed CLI configuration.

Open Terminal A and run:

bash
claude --version
codex --version
python3 --version
git --version
curl --version
jq --version
codex login status

Expected: each command is found, and the last command reports Logged in using ChatGPT. If it does not, run codex login, finish the browser login, then check again. If Claude needs sign-in, launch claude, finish sign-in, and exit back to the shell. No API key is needed for this walkthrough.

  • Install this PR's revision. Paste the following into the same Terminal A. It makes two temporary Git repositories and installs Quest in each. The local commits only establish a clean starting point; nothing is pushed.
set -e
quest176_sha=ddf94260f2671adfa8e4ec556b275a5939538ad7
quest176_lab=$(mktemp -d /tmp/quest-176-review.XXXXXX)
mkdir -p "$quest176_lab/artifacts"

test "$(git ls-remote https://github.com/KjellKod/quest.git refs/heads/quest/codex-cli-dispatch | cut -f1)" = "$quest176_sha"
for host in Claude Codex; do
  fixture="$quest176_lab/$host consumer"
  mkdir -p "$fixture"
  git -C "$fixture" init -b quest/acceptance
  curl -fsSL "https://raw.githubusercontent.com/KjellKod/quest/$quest176_sha/scripts/quest_installer.sh" -o "$fixture/quest_installer.sh"
  (
    cd "$fixture"
    bash ./quest_installer.sh --branch quest/codex-cli-dispatch --force
    test "$(cat .quest-version)" = "$quest176_sha"
    printf '\n__pycache__/\n*.pyc\n' >> .git/info/exclude
    git add .
    git -c user.name='Quest test' -c user.email='quest-test@example.invalid' commit -m 'Prepare Quest test folder'
    test -z "$(git status --porcelain)"
  )
done
test "$(git ls-remote https://github.com/KjellKod/quest.git refs/heads/quest/codex-cli-dispatch | cut -f1)" = "$quest176_sha"
printf '\nPASS: installed %s\nTest folders: %s\n' "$quest176_sha" "$quest176_lab"
printf "\nClaude launch command:\ncd '%s/Claude consumer' && claude --model claude-opus-5\n" "$quest176_lab"
printf "\nCodex launch command:\ncd '%s/Codex consumer' && codex --model gpt-6-astra -c 'model_reasoning_effort=\"medium\"'\n" "$quest176_lab"

Expected: the final PASS: installed ... line contains the revision above, followed by two launch commands with real folder paths. Keep Terminal A open, because later checks use its quest176_lab variable.

If setup stops: do not continue. If the PR branch has moved, the revision guard intentionally stops the test so a different version cannot be mistaken for this one. Record the failed command and error. Do not switch to main to get past it.

2. Ask Claude to use Codex once

  • Action: open Terminal B. Copy the Claude launch command printed by setup into Terminal B. Accept the normal workspace trust prompt for this temporary folder. Then paste this into the Claude conversation, not the shell:
/gpt Ask Codex to calculate 17 times 19 and return the answer. Use this
folder's installed Quest gpt skill and scripts/quest_codex_runner.py,
cached ChatGPT login, gpt-6-astra, medium effort, and read-only sandbox.
Store the runner evidence under .quest/gpt-proof. Do not change project
files. Afterward show the runner command, receipt path, Codex session ID,
result_kind and cleanup result. Do not answer the calculation yourself
instead of calling Codex.

Expected: answer 323, plus a successful tool call to scripts/quest_codex_runner.py task. Its receipt should say result_kind: complete, auth_mode: cached, auth_kind: chatgpt, and cleanup: complete.

Fail: Claude only answers 323, uses a Codex MCP tool, cannot find /gpt, or reports a failed runner. The answer alone does not prove delegation. Save the error/receipt; do not install a shim or change authentication to make this test pass.

3. Run a Claude-led Quest, with a real approval pause

  • Action: in the same Claude conversation in Terminal B, paste:
/quest Use full workflow mode in this existing test branch. Create only
hello-world.txt containing exactly hello world followed by one newline
(12 bytes). Use installed Quest skills and saved defaults: Astra medium
for Codex roles, Claude Opus 5 for Claude roles, cached ChatGPT login.
Complete planning, both plan reviews, arbiter and walkthrough. Explain
what the plan will do, then STOP for my explicit Build approval. The file
must not exist before approval. After approval, build and complete both
code reviews and any required fixes. Claude must launch Codex builder and
reviewer roles through the installed Quest runner. Use the true saved
review counter, including --iter 0 for the first code review. Keep evidence
under .quest. Do not commit, push, merge, broaden permissions or switch
models/authentication to work around a failure.

If Quest asks which models to use, keep the defaults listed in the prompt.

Expected pause: Claude presents a reviewed plan and asks whether to Build. It may create planning records under .quest; it must not create hello-world.txt yet.

In Terminal A, select the Claude target:

quest176_target="$quest176_lab/Claude consumer/hello-world.txt"

Run the Before approval check below. Only if it passes, send this in the Claude conversation:

Build approved for this hello-world test Quest.

Expected after approval: Claude has Codex create the file, then runs two code reviews. Run the After Build check below when the file has been created. If the assistant pauses for that check, paste the actual PASS output back into the conversation so review can continue.

4. Repeat with Codex leading

  • Action: exit the Claude session in Terminal B, then run the Codex launch command printed by setup. In the Codex conversation, paste:
$quest Use full workflow mode in this existing test branch. Create only
hello-world.txt containing exactly hello world followed by one newline
(12 bytes). Use installed Quest skills and saved defaults: Astra medium
for Codex roles, Claude Opus 5 for Claude roles, cached ChatGPT login.
Complete planning, both plan reviews, arbiter and walkthrough. Explain
what the plan will do, then STOP for my explicit Build approval. The file
must not exist before approval. After approval, build and complete both
code reviews and any required fixes. Codex roles must use native Codex
subagents, never Codex MCP or nested codex exec. If native delegation or
the required settings cannot be honored, report the blocker. Keep evidence
under .quest. Do not commit, push, merge, broaden permissions or switch
models/authentication to work around a failure.

In Terminal A, select the Codex target:

quest176_target="$quest176_lab/Codex consumer/hello-world.txt"

At Codex's approval pause, run Before approval. Then send Build approved for this hello-world test Quest. in the Codex conversation. Run After Build when the file is created, and let both code reviews finish.

Expected: the same exact file, but the Codex roles appear as native helper-agent calls. A terminal command that starts another Codex process is not a substitute for native delegation.

File checks, use these for each Quest

  • Before approval, paste into Terminal A after selecting the correct target above:
python3 - "$quest176_target" <<'PY'
import os, sys
assert not os.path.lexists(sys.argv[1]), "FAIL: target already exists before approval"
print("PASS: target absent before Build approval")
PY

Expected: PASS: target absent before Build approval. If the file already exists, stop that test and save the evidence. Deleting it afterward does not restore proof of the approval gate.

  • After Build, paste into Terminal A:
python3 - "$quest176_target" <<'PY'
from pathlib import Path
import sys
p = Path(sys.argv[1])
assert p.is_file() and not p.is_symlink(), "FAIL: expected a regular file"
data = p.read_bytes()
assert data == b"hello world\n", f"FAIL: unexpected bytes: {data!r}"
print("PASS: exact hello world plus one newline, 12 bytes")
PY

Expected: PASS: exact hello world plus one newline, 12 bytes. This checks the newline too; seeing the words in an editor is insufficient.

Finish: save a result that someone else can verify

  • After each Quest finishes, paste this into its assistant conversation:
Write a short validation report under this Quest's .quest folder. Use PASS,
FAIL or NOT VERIFIED for each check. Include the installed revision, the
recorded Build approval and file-creation order, exact-byte check, both
review results, and paths to validated handoff files (the structured role
completion records). Separate checks I actually ran from automated checks.

Prove the dispatch route from the real tool/session records. For Claude,
show the builder/reviewer runner commands, first review iteration 0, and
matching Codex child sessions. For Codex, show native helper-agent calls
and parent/child session IDs; inspect the full relevant trace for Codex MCP
or nested CLI self-calls. Check saved versus effective model/effort using
runtime metadata, not a model's claim about its identity. Include auth mode
and cleanup receipts without secrets. Mark any unavailable trace or check
NOT VERIFIED. Show the report path and a short result table.

Human sign-off: both files pass, both approval pauses happened before file creation, both pairs of code reviews finished, and the report points to real dispatch evidence. An assistant saying “all good” without those records is insufficient. No need to understand every JSON field or rerun CI yourself. If evidence is unclear, leave that row NOT VERIFIED and attach the report for a maintainer to inspect.

When blocked: record the route being tested, last command/tool error, installed revision and report path. Do not paste credentials. Keep the temporary folders until the review is finished.

Advanced checks for maintainers: external installations, upgrades, API keys and automated regressions

The walkthrough above uses the ordinary inside-repository installation. The following checks cover additional layouts and authentication. They are separate from the three human checks above. API-key inference remains NOT VERIFIED unless someone with an authorized key observes a successful real call.

These commands use the quest176_lab and quest176_sha variables in Terminal A. Refer to setup above for the pinned installer download and revision guard.

External installation, non-Git folders and upgrades

In Terminal A, install Quest outside a repository. This folder holds the helper; a different folder will receive the output file:

mkdir -p "$quest176_lab/Outside installation"
curl -fsSL "https://raw.githubusercontent.com/KjellKod/quest/$quest176_sha/scripts/quest_installer.sh" -o "$quest176_lab/Outside installation/quest_installer.sh"
(
  cd "$quest176_lab/Outside installation"
  bash ./quest_installer.sh --branch quest/codex-cli-dispatch --force
  test "$(cat .quest-version)" = "$quest176_sha"
)

Then create a separate target repository and run the installed helper:

mkdir -p "$quest176_lab/Separate target" "$quest176_lab/Non Git target"
git -C "$quest176_lab/Separate target" init -b quest/acceptance
printf 'Create hello-world.txt containing exactly hello world followed by one LF.\n' |
  python3 "$quest176_lab/Outside installation/scripts/quest_codex_runner.py" task \
    --cwd "$quest176_lab/Separate target" \
    --output-dir "$quest176_lab/artifacts/external-git" \
    --auth cached --model gpt-6-astra --effort medium \
    --sandbox workspace-write --timeout 1800

Expected: current completed receipt, file only in the separate target, artifacts outside both roots, exact byte check passes. Repeat against Non Git target: without --allow-non-git require nonzero exit and absent target; with that flag and a fresh output directory require successful inference and exact bytes. Verify actual runtime cwd/sandbox/writable artifact root, not only argv.

For the legacy upgrade baseline, the installer accepts a branch name, not a commit SHA. The observed baseline came from main while main pointed at 319494e. Reproduce it only while this guard passes:

quest176_baseline=319494ec43e53906836e8054aebedabeca1f8314
test "$(git ls-remote https://github.com/KjellKod/quest.git refs/heads/main | cut -f1)" = "$quest176_baseline"
mkdir -p "$quest176_lab/Legacy pristine" "$quest176_lab/Upgrade home"
git -C "$quest176_lab/Legacy pristine" init -b quest/upgrade-acceptance
curl -fsSL "https://raw.githubusercontent.com/KjellKod/quest/$quest176_baseline/scripts/quest_installer.sh" \
  -o "$quest176_lab/Legacy pristine/quest_installer.sh"
(
  cd "$quest176_lab/Legacy pristine"
  HOME="$quest176_lab/Upgrade home" PATH=/usr/bin:/bin:/usr/sbin:/sbin \
    bash ./quest_installer.sh --branch main --force
  test "$(cat .quest-version)" = "$quest176_baseline"
)
cp -R "$quest176_lab/Legacy pristine" "$quest176_lab/Legacy custom"

If main moved, stop this legacy reproduction. Do not label a newer installation as the old baseline. Existing recorded baseline receipts remain historical evidence.

In Legacy custom, add consumer-owned.txt and an unrelated MCP entry to .opencode/opencode.json; record SHA256 hashes of both before upgrading. Keep Legacy pristine unchanged. In each directory, download the candidate installer from the pinned candidate URL in step 1, run it with --branch quest/codex-cli-dispatch --force using the isolated Upgrade home/PATH above, then require .quest-version to equal the candidate SHA.

Expected: pristine managed OpenCode Codex MCP removed; customized configuration and consumer-owned.txt retain their exact hashes, with targeted migration guidance instead of blanket deletion. Actual before/after evidence: distribution/final-legacy-preservation-before.json and distribution/final-preservation-after.json.

These integration checks do not substitute for the Claude-led and Codex-led Quests above.

API-key inference without cached login, conditional

Only run when an authorized CODEX_API_KEY or OPENAI_API_KEY is already available. Never paste/log it. In Terminal A, use the installed Claude test folder. This is a direct runner check, separate from the Claude conversation:

cd "$quest176_lab/Claude consumer"
quest176_api_home=$(mktemp -d /tmp/quest-176-api-home.XXXXXX)
printf 'Create api-probe.txt containing exactly hello world followed by one LF.\n' |
  CODEX_HOME="$quest176_api_home" python3 scripts/quest_codex_runner.py task \
    --cwd "$PWD" --output-dir "$quest176_lab/artifacts/api-key" \
    --auth api-key --model gpt-6-astra --effort medium \
    --sandbox workspace-write --timeout 1800

PASS requires: real successful inference, auth_mode=api-key, auth_kind=api-key, matching runtime model/effort and exact api-probe.txt bytes, despite no prior CLI login in that new home. A dummy-key endpoint or readiness probe is not successful OpenAI inference. No authorized key means NOT VERIFIED.

Automated checks, optional local reproduction of CI

python3 -m pip install -e '.[dev]'
python3 -m pytest tests/ -q
python3 -m black --check .
bash tests/test-quest-preflight.sh
bash tests/test-quest-runtime.sh
bash tests/test-quest-orchestration.sh
bash scripts/quest_validate-manifest.sh
bash scripts/quest_validate-quest-config.sh
bash scripts/quest_validate-handoff-contracts.sh
python3 scripts/quest_sync_model_defaults.py --check

The five schema regression cases require the AJV CLI with draft2020 support; they skip if unavailable. All five passed locally on this head.

The two tests in tests/integration/test_codex_mcp_overrides.py require a real installed Codex CLI, otherwise they skip. Local PASS on Codex 0.154.0 is recorded separately. Deterministic tests cover credential isolation, stdin/argv propagation, malformed/stale/missing outputs, process-tree timeout/cancellation, retry rules and unavailable/mismatched native controls.

Notes

No merge. Both full host workflows have run with explicit Build approvals. The failed initial-review path was fixed, then passed a separate actual Claude-launched installed replay. Retained older live proof and later source/installed equivalence are explicitly separated. Live API-key inference remains conditional and unverified.

Codex acceptance planning used f649beb; the unchanged reviewed plan was built and code-reviewed on then-current candidate 0ec404d. Executable dispatch arguments and tool names were inspectable throughout; some collaboration prompt/message bodies were encrypted. Builder and native reviewer bootstrap read-only commands exceeded their declared role lists. External-symlink journal discovery failed; a manual canonical journal and installed --skip-journal archive completed the workflow. These limits are recorded, without claiming strict role allowlist compliance or global process cleanup.

Normal host sessions added Codex project registrations; structured comparison confirms MCP settings and model/effort unchanged. Consumer upgrade sentinels remain byte-identical. Existing reverse-direction Claude probe respawn behavior is recorded separately, without expanding this migration.

     ▐▛███▜▌
    ▝▜█████▛▘
      ▘▘ ▝▝
Quest/Co-Authored by
Co-Authored-By: GPT-6 Astra <noreply@openai.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
in collaboration with KjellKod

Review in cubic

KjellKod and others added 2 commits September 12, 2026 14:56
Use explicit cached or API-key authentication for Claude-to-Codex tasks
and Quest roles, preserving native Codex delegation and installed paths.
Correct installer/preflight guidance and validate artifacts and cleanup.

Refs #176. Required installed live acceptance is tracked separately.

Quest/Co-Authored by
Co-Authored-By: GPT-6 Astra <noreply@openai.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Use a TOML table value supported by the actual CLI override parser.
Preserve unrelated MCP entries and cover dashed and dotted names with
real configuration parsing regressions.

Refs #176

Co-Authored-By: GPT-6 Astra <noreply@openai.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: KjellKod <KjellKod@users.noreply.github.com>
@KjellKod
KjellKod marked this pull request as ready for review September 12, 2026 22:08
@KjellKod
KjellKod deployed to codex-ci-review September 12, 2026 22:08 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3 issues found across 34 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="scripts/quest_preflight.sh">

<violation number="1" location="scripts/quest_preflight.sh:318">
P2: When the codex runner probe fails (timeout, crash, or non-JSON output), the fallback discards the runner's captured stderr and emits a generic "Reinstall Quest and rerun preflight" warning with codex_auth_reason set to "not_checked". In this narrow failure mode the real diagnostic is the most valuable output, but it is never surfaced. Print the captured stderr into the warning array so the operator can act on the actual failure instead of being told to reinstall.</violation>
</file>

<file name="scripts/quest_runtime/codex_runner.py">

<violation number="1" location="scripts/quest_runtime/codex_runner.py:641">
P2: In `_validate_handoff`, the codex handoff is validated against the `next` and `artifacts` keys, but the canonical handoff contract in `.ai/schemas/handoff.schema.json` defines `next_role` (required) and `artifacts_written` with `additionalProperties: false`. If the Codex agent produces a schema-compliant handoff, every role output is rejected as `malformed_output` and the retry ladder loops forever; if it produces the `workflow.md` format, the schema still rejects the file. The new validator should read the same field names the agent is instructed to write (`next`/`artifacts` per `.skills/quest/delegation/workflow.md` line 455) or the schema's names, and the repo's two handoff contracts should be reconciled before this validation is relied on.</violation>

<violation number="2" location="scripts/quest_runtime/codex_runner.py:761">
P2: Codex context-health entries are written without `model=` and `effort=` fields, even though the documented Codex log contract for this workflow requires those fields (with legacy compatibility when present). `append_context_health_log` in claude_runner.py has no `model`/`effort` parameters, so the new Codex role path cannot emit them. Add optional `model`/`effort` params to the helper and pass the saved `model` and `effort` in this call so Codex lines carry the effective settings.</violation>
</file>

Tip: cubic can generate docs of your entire codebase and keep them up to date. Try it here.

Re-trigger cubic

Comment thread scripts/quest_preflight.sh Outdated
Comment thread scripts/quest_runtime/codex_runner.py Outdated
Comment thread scripts/quest_runtime/codex_runner.py Outdated
Comment thread scripts/quest_codex_runner.py
Comment thread scripts/quest_codex_runner.py Outdated
Comment thread tests/test-quest-preflight.sh
Comment thread .ai/allowlist.json
Comment thread tests/integration/test_codex_mcp_overrides.py
Comment thread scripts/quest_installer.sh Outdated
Comment thread tests/unit/test_codex_runner.py
Reject invalid timeout and model inputs before dispatch. Preserve known execution
and authentication metadata when receipt or role logging fails. Return structured
preflight diagnostics without exposing captured secrets, keep optional installer
scans nonfatal, and align authentication examples and validation schema.

Quest/Co-Authored by
Co-Authored-By: GPT-6 Astra <noreply@openai.com>
in collaboration with KjellKod
@KjellKod
KjellKod deployed to codex-ci-review September 12, 2026 22:53 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 12 files (changes from recent commits).

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread scripts/quest_preflight.sh
Keep the runner aligned with Quest fix_iteration, which starts at zero.
Cover both reviewer slots with fresh artifact, receipt and logging checks.

Quest/Co-Authored by
Co-Authored-By: GPT-6 Astra <noreply@openai.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
in collaboration with KjellKod
@KjellKod
KjellKod deployed to codex-ci-review September 13, 2026 05:10 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 3 files (changes from recent commits).

Requires human review: Auto-approval blocked by 7 unresolved issues from previous reviews.

Re-trigger cubic

The Codex plan-reviewer example prompt allowed NEXT: null when blocked,
but the handoff contract and runner require next: arbiter, so an honest
blocked handoff was rejected as malformed and retried once for nothing.

The installer overwrites files with mv and only restores the executable
bit for EXECUTABLE_FILES. session-start.sh was never listed, so any
update that rewrote the hook left it non-executable and Claude's
SessionStart hook failed with permission denied. Add it to the list and
assert every manifest hook under .claude/hooks is covered.

Both surfaced by Codex review on a consumer install of this branch.

Quest/Co-Authored by
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
@KjellKod
KjellKod deployed to codex-ci-review September 21, 2026 02:44 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files (changes from recent commits).

Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread tests/unit/test_source_python_formatting.py
@KjellKod

Copy link
Copy Markdown
Owner Author

This works, I haven't done all the validation yet, but I've been running this, and it works fine. I will do PR-shepherd. And all the validation data. This is not a high priority right now, but if people want to use this, they can.

@KjellKod
KjellKod deployed to codex-ci-review September 22, 2026 13:55 — with GitHub Actions Active
@KjellKod
KjellKod deployed to codex-ci-review September 22, 2026 13:58 — with GitHub Actions Active
@KjellKod
KjellKod merged commit 2cf4d58 into main Sep 22, 2026
7 checks passed
@KjellKod
KjellKod deleted the quest/codex-cli-dispatch branch September 22, 2026 14:01

This branch was successfully deployed

1 active deployment
codex-ci-review — 510f204f Deployed Sep 22, 2026 by KjellKod via codex-review #731
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant