Repository navigation
Conversation
Every Quest role ran at a top-tier model with effort left to each runtime's default, and the Claude side had no effort control at all. Two of the four dispatch paths silently ignored the setting: the Codex-led Claude runner built argv with --model only, and Claude-led Claude roles run as native Task() calls, which read effort from subagent frontmatter rather than orchestration.json. Add an `effort` map to the allowlist, parallel to `models`, resolved per role on whichever runtime that role's model selects. The generator now writes it to all four surfaces, including `effort:` frontmatter in .claude/agents, so the allowlist stays the single source of truth. Quests written before the map keep resolving through the legacy codex_reasoning_effort scalar, so resume never re-tiers work already in flight. Effort is tiered rather than flat: high for planning, building, code review and both arbiters, medium for plan review and fixing. Effort pays most on coding and long-horizon work, and the review roles are diff-scoped, so depth there is cheap next to the fix loop it avoids. Also move the four Claude roles to claude-opus-5-5, switch review_mode to auto so small diffs stop drawing full-context reviews, and cap max_plan_iterations at 3. Quest/Co-Authored by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Measurement, not assumption: varying only how effort is supplied, the `claude --effort` flag separates output tokens by roughly 10x (low 111/99/84 vs max 1108/1017/1209), while subagent-file `effort:` frontmatter showed no separation (low 245/282/214 vs max 258/243/239). An invalid value in a file also raises no warning, where the flag warns and falls back. Keep generating the frontmatter — it is the documented control, the only one available on that path, and free to emit — but stop stating it as working. Route effort-critical roles through the runner instead. Quest/Co-Authored by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
KjellKod
commented
Sep 24, 2026
KjellKod
commented
Sep 24, 2026
KjellKod
commented
Sep 24, 2026
KjellKod
commented
Sep 24, 2026
KjellKod
commented
Sep 24, 2026
Field observation beats the general guidance I based the tiers on: Astra and Opus 5.5 both start gold-plating at high effort, inventing work rather than doing the asked work better. The "coding responds to effort" evidence comes from closed-scope benchmarks where extra deliberation can only improve the answer; an open-ended build task has unbounded surface for it to spend on. The pipeline shape reinforces this. Two reviewers, an arbiter and a fixer catch under-doing well. They catch gold-plating poorly, since reviewers rarely argue for deleting work and the arbiter filters noise, not scope. The asymmetry favors restraint in every generative role. Flatten the map to medium everywhere and pin the builder to gpt-5.6-sol, which this repo's own field notes describe as the disciplined, compliant builder. The per-role map stays as the control surface for deviating later; it just no longer deviates by default. Quest/Co-Authored by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
An Astra review of the branch found six issues; four were confirmed by repro and are fixed here. Legacy quests were re-tiered on resume. Snapshot migration passed no effort map to the writer, which then filled it with today's defaults, so a quest started with codex_reasoning_effort=high resolved to medium. The writer now omits the key entirely when the caller has no effort policy, and only the new-quest path synthesizes defaults. The legacy scalar is now Codex-scoped, as its name always implied. Before the per-role map, the Claude runner never read it, so applying it to Claude roles both re-tiered legacy work and turned a quest pinned to the Codex-only `ultra` into a hard Claude dispatch failure. Both Tier B permission retries dropped effort: every other argument was forwarded to the recursive call except this one, so a pinned role retried at the default. That is precisely the run where the pin matters. A malformed effort map silently fell back to the legacy scalar, turning a config error into a plausible wrong answer; it now raises. The state validator also accepted unknown role keys, so a typo passed validation and was ignored. The frontmatter generator now strips quoted spellings of the key, which YAML's last-key-wins would otherwise honor over the generated one. Regression tests cover each case. The retry test drives the real Tier B path and was verified to fail with the fix removed. Quest/Co-Authored by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: GPT-6 Astra <noreply@openai.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Astra re-reviewed its own five fixes adversarially and found more: snapshot migration still flattened PARTIAL effort maps to defaults, an explicit `"effort": null` resolved as absent, unreadable config was treated as missing, and rejecting an effort level truncated existing artifacts because validation ran after artifact preparation. It also added a bounded guard for finding 4 rather than converting native dispatch to the runner, and fixed a pre-existing solo-mode migration bug it hit while testing. Reviewing that work turned up two over-corrections, both fixed here. The native guard blocked any role the quest does not pin. Since the generated frontmatter always names a level and an unpinned role never can, that walled off every legacy quest on resume — the exact case the previous two commits existed to protect. An unpinned role now passes; only a genuine disagreement blocks. The guard also rejected any agent file with nested YAML, so adding a valid `hooks:` or `mcpServers:` block would have hard-blocked dispatch. Nested blocks are now skipped, while a stray `effort:` at any depth still fails closed as ambiguous. Validating effort early also started raising out of run_claude_role, replacing the JSON envelope orchestrators parse with a traceback. The ordering fix is kept; the structured invocation_error is restored. Quest/Co-Authored by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: GPT-6 Astra <noreply@openai.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
KjellKod
marked this pull request as ready for review
September 24, 2026 23:03
There was a problem hiding this comment.
All reported issues were addressed across 24 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
Quest/Co-Authored by Co-Authored-By: GPT-6 <noreply@openai.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
There was a problem hiding this comment.
All reported issues were addressed across 10 files (changes from recent commits).
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
Quest/Co-Authored by Co-Authored-By: GPT-6 <noreply@openai.com> in collaboration with KjellKod <kjell.hedstrom@gmail.com>
There was a problem hiding this comment.
0 issues found across 1 file (changes from recent commits).
Requires human review: Wires per-role effort through dispatch, moves Claude roles to Opus 5.5, and changes review_mode/plan-iteration defaults. Human review required: native Task() effort is explicitly unverified; model/effort defaults are spend/quality tradeoffs.
Re-trigger cubic
13 of 15 tasks
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
⭐ Why this matters
Quest needs a consistent place to configure reasoning effort across its role dispatch paths. This change makes effort a per-role allowlist setting and carries it through saved quest configuration and runtime dispatch.
Summary
Adds per-role effort, defaults every role to
medium, moves the four Claude roles toclaude-opus-5-5, and pins the builder togpt-5.6-sol.mediumis the selected default for scope restraint. The validation below does not establish lower spend or better task quality.Warning
Hold merge: customized-install upgrade regression reproduced. The installer preserves a customized agent file without
effort:, while a new quest savesmedium. The native Claude guard then blocks dispatch. This requires a supported upgrade path that preserves customizations and saved settings.Changes
effortmap and schema validation; changesreview_modefromfulltoautoandgates.max_plan_iterationsfrom4to3.quest_claude_runner.pyand forwards it through background-agent and legacy bridge commands. The active workflow uses background agents.effort:frontmatter and checks it against saved settings before dispatch.Validation completed
Completed agentically on 2026-09-30, at commit
4e4aa756bcf6f03f650a9ddf845d6f3719e9fbcd, on macOS with Claude Code 2.1.284 and Codex CLI 0.157.1. Reviewers do not need to repeat these checks.PYTHONPATH=scripts python3 -m pytest tests/ -q: 1,389 passed.bash tests/test-quest-orchestration.sh: 36/36 passed;bash tests/test-quest-runtime.sh: 72/72 passed;bash tests/test-validate-quest-state.sh: 72/72 passed.python3 -m black --check .andpython3 scripts/quest_sync_model_defaults.py --check: passed. A disposable fixture with changed arbiter effort correctly reported stale runtime defaults and agent frontmatter.modelsand unsupported Claudeultraeffort with structured errors, preserving pre-existing output artifacts.claude-opus-5-5atmedium.Agentdispatch returned the expected fixture marker. With the parent atmedium, children configured through frontmatter recordedlowandhighrespectively in their own assistant transcripts. This validates native dispatch on the tested version, not merely--agentsession activation or response length.gpt-5.6-sol/mediumreturned the exact fixture marker.gpt-5.6-sol/medium. The runner accepted the synthetic handoff and output artifacts; their contents were independently verified.Evidence and limits
Claude background session
5c811073-aed3-4fbe-9a05-dce918a0d8e8completed the review probe. Native low/high comparison session:b54b0b30-bb44-4385-994e-f0742ab553af. Codex runner session:01a0f4a1-4f7e-78a2-865f-0426d477ac7e.The Codex runner receipt leaves
effective_modelandeffective_effortnull. Requested settings and successful artifacts are verified; independently reported effective Codex settings are not. Claude transcript effort is runtime metadata, not proof of reasoning quality or cost savings.All live Claude checks in this validation used background agents. The legacy API bridge was not called and is not a pending human validation requirement for this workflow. Its flag forwarding has automated coverage. No full feature quest was built or merged.
The earlier output-length experiments using
--agenttested a different activation path. They do not supersede the direct nativeAgenttranscript evidence above.Remaining work before merge
Fix and regression-test the customized-install upgrade path. Reproduction: preserve a locally customized pre-change
.claude/agents/arbiter.mdduring an installer update, create a new quest, then callvalidate_native_claude_effort. It currently raises:Relevant code:
scripts/quest_installer.sh(install_copy_as_is_file) andscripts/quest_runtime/orchestration.py(write_default_from_allowlist,validate_native_claude_effort). This is implementation work, not a request for the reviewer to repeat the test suites.Human review should focus on the intended policy choices:
mediumeffort,autoreview mode, and the three-iteration planning limit. Do not infer their quality/cost tradeoffs from these smoke tests.▐▛███▜▌ ▝▜█████▛▘ ▘▘ ▝▝ Quest/Co-Authored by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: GPT-6 <noreply@openai.com> in collaboration with KjellKod