Skip to content

Add per-role reasoning effort, default to medium, move Claude roles to Opus 5.5 - #179

Closed
KjellKod wants to merge 7 commits into
mainfrom
worktree-quest-effort-and-opus55
Closed

KjellKod wants to merge 7 commits into
mainfrom
worktree-quest-effort-and-opus55

Conversation

@KjellKod

@KjellKod KjellKod commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

⭐ Why this matters

Quest needs a consistent place to configure reasoning effort across its role dispatch paths. This change makes effort a per-role allowlist setting and carries it through saved quest configuration and runtime dispatch.


Summary

Adds per-role effort, defaults every role to medium, moves the four Claude roles to claude-opus-5-5, and pins the builder to gpt-5.6-sol.

  • Existing quests preserve their saved effort settings and the Codex-scoped legacy fallback.
  • medium is the selected default for scope restraint. The validation below does not establish lower spend or better task quality.

Warning

Hold merge: customized-install upgrade regression reproduced. The installer preserves a customized agent file without effort:, while a new quest saves medium. The native Claude guard then blocks dispatch. This requires a supported upgrade path that preserves customizations and saved settings.

Changes

  • Configuration: adds the per-role effort map and schema validation; changes review_mode from full to auto and gates.max_plan_iterations from 4 to 3.
  • Claude dispatch: resolves saved role effort in quest_claude_runner.py and forwards it through background-agent and legacy bridge commands. The active workflow uses background agents.
  • Codex dispatch: role mode reads saved per-role effort, with the legacy scalar as fallback.
  • Native Claude dispatch: generates agent effort: frontmatter and checks it against saved settings before dispatch.
  • Compatibility and validation: preserves legacy effort during resume, rejects malformed configuration, preserves nested frontmatter content, and adds regression coverage.

Validation completed

Completed agentically on 2026-09-30, at commit 4e4aa756bcf6f03f650a9ddf845d6f3719e9fbcd, on macOS with Claude Code 2.1.284 and Codex CLI 0.157.1. Reviewers do not need to repeat these checks.

  • PYTHONPATH=scripts python3 -m pytest tests/ -q: 1,389 passed.
  • bash tests/test-quest-orchestration.sh: 36/36 passed; bash tests/test-quest-runtime.sh: 72/72 passed; bash tests/test-validate-quest-state.sh: 72/72 passed.
  • python3 -m black --check . and python3 scripts/quest_sync_model_defaults.py --check: passed. A disposable fixture with changed arbiter effort correctly reported stale runtime defaults and agent frontmatter.
  • Actual runner CLI rejected malformed saved models and unsupported Claude ultra effort with structured errors, preserving pre-existing output artifacts.
  • Real Codex-led Claude background dispatch resolved saved effort without a CLI effort override, produced its review and handoff artifacts, and cleaned up without fallback. Assistant transcript records claude-opus-5-5 at medium.
  • Real native Claude Agent dispatch returned the expected fixture marker. With the parent at medium, children configured through frontmatter recorded low and high respectively in their own assistant transcripts. This validates native dispatch on the tested version, not merely --agent session activation or response length.
  • Native Codex dispatch with requested gpt-5.6-sol / medium returned the exact fixture marker.
  • A real Claude agent invoked the saved-config Codex builder runner using cached ChatGPT authentication and requested gpt-5.6-sol / medium. The runner accepted the synthetic handoff and output artifacts; their contents were independently verified.
  • Independent Claude review completed; its customized-install upgrade concern was independently reproduced using the actual installer function, prior shipped-file checksum, new-quest writer, and native effort guard. Only the network fetch was replaced with the local upstream fixture.

Evidence and limits

Claude background session 5c811073-aed3-4fbe-9a05-dce918a0d8e8 completed the review probe. Native low/high comparison session: b54b0b30-bb44-4385-994e-f0742ab553af. Codex runner session: 01a0f4a1-4f7e-78a2-865f-0426d477ac7e.

The Codex runner receipt leaves effective_model and effective_effort null. Requested settings and successful artifacts are verified; independently reported effective Codex settings are not. Claude transcript effort is runtime metadata, not proof of reasoning quality or cost savings.

All live Claude checks in this validation used background agents. The legacy API bridge was not called and is not a pending human validation requirement for this workflow. Its flag forwarding has automated coverage. No full feature quest was built or merged.

The earlier output-length experiments using --agent tested a different activation path. They do not supersede the direct native Agent transcript evidence above.

Remaining work before merge

Fix and regression-test the customized-install upgrade path. Reproduction: preserve a locally customized pre-change .claude/agents/arbiter.md during an installer update, create a new quest, then call validate_native_claude_effort. It currently raises:

Native Claude effort mismatch for arbiter: saved='medium', frontmatter=None. Stop before Task dispatch.

Relevant code: scripts/quest_installer.sh (install_copy_as_is_file) and scripts/quest_runtime/orchestration.py (write_default_from_allowlist, validate_native_claude_effort). This is implementation work, not a request for the reviewer to repeat the test suites.

Human review should focus on the intended policy choices: medium effort, auto review mode, and the three-iteration planning limit. Do not infer their quality/cost tradeoffs from these smoke tests.

     ▐▛███▜▌
    ▝▜█████▛▘
      ▘▘ ▝▝
Quest/Co-Authored by
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: GPT-6 <noreply@openai.com>
in collaboration with KjellKod

KjellKod and others added 2 commits September 22, 2026 17:00
Every Quest role ran at a top-tier model with effort left to each
runtime's default, and the Claude side had no effort control at all.
Two of the four dispatch paths silently ignored the setting: the
Codex-led Claude runner built argv with --model only, and Claude-led
Claude roles run as native Task() calls, which read effort from
subagent frontmatter rather than orchestration.json.

Add an `effort` map to the allowlist, parallel to `models`, resolved
per role on whichever runtime that role's model selects. The generator
now writes it to all four surfaces, including `effort:` frontmatter in
.claude/agents, so the allowlist stays the single source of truth.
Quests written before the map keep resolving through the legacy
codex_reasoning_effort scalar, so resume never re-tiers work already
in flight.

Effort is tiered rather than flat: high for planning, building, code
review and both arbiters, medium for plan review and fixing. Effort
pays most on coding and long-horizon work, and the review roles are
diff-scoped, so depth there is cheap next to the fix loop it avoids.

Also move the four Claude roles to claude-opus-5-5, switch review_mode
to auto so small diffs stop drawing full-context reviews, and cap
max_plan_iterations at 3.

Quest/Co-Authored by
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Measurement, not assumption: varying only how effort is supplied, the
`claude --effort` flag separates output tokens by roughly 10x
(low 111/99/84 vs max 1108/1017/1209), while subagent-file `effort:`
frontmatter showed no separation (low 245/282/214 vs max 258/243/239).
An invalid value in a file also raises no warning, where the flag warns
and falls back.

Keep generating the frontmatter — it is the documented control, the only
one available on that path, and free to emit — but stop stating it as
working. Route effort-critical roles through the runner instead.

Quest/Co-Authored by
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Comment thread .claude/agents/builder.md Outdated
Comment thread .claude/agents/code-reviewer.md Outdated
Comment thread .ai/allowlist.json Outdated
Comment thread .claude/agents/arbiter.md Outdated
Comment thread .claude/agents/review-arbiter.md Outdated
Field observation beats the general guidance I based the tiers on:
Astra and Opus 5.5 both start gold-plating at high effort, inventing
work rather than doing the asked work better. The "coding responds to
effort" evidence comes from closed-scope benchmarks where extra
deliberation can only improve the answer; an open-ended build task has
unbounded surface for it to spend on.

The pipeline shape reinforces this. Two reviewers, an arbiter and a
fixer catch under-doing well. They catch gold-plating poorly, since
reviewers rarely argue for deleting work and the arbiter filters noise,
not scope. The asymmetry favors restraint in every generative role.

Flatten the map to medium everywhere and pin the builder to
gpt-5.6-sol, which this repo's own field notes describe as the
disciplined, compliant builder. The per-role map stays as the control
surface for deviating later; it just no longer deviates by default.

Quest/Co-Authored by
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
@KjellKod KjellKod changed the title Pin reasoning effort per role and move Claude roles to Opus 5.5 Add per-role reasoning effort, default to medium, move Claude roles to Opus 5.5 Sep 24, 2026
KjellKod and others added 2 commits September 23, 2026 21:33
An Astra review of the branch found six issues; four were confirmed by
repro and are fixed here.

Legacy quests were re-tiered on resume. Snapshot migration passed no
effort map to the writer, which then filled it with today's defaults,
so a quest started with codex_reasoning_effort=high resolved to medium.
The writer now omits the key entirely when the caller has no effort
policy, and only the new-quest path synthesizes defaults.

The legacy scalar is now Codex-scoped, as its name always implied.
Before the per-role map, the Claude runner never read it, so applying
it to Claude roles both re-tiered legacy work and turned a quest
pinned to the Codex-only `ultra` into a hard Claude dispatch failure.

Both Tier B permission retries dropped effort: every other argument was
forwarded to the recursive call except this one, so a pinned role
retried at the default. That is precisely the run where the pin matters.

A malformed effort map silently fell back to the legacy scalar, turning
a config error into a plausible wrong answer; it now raises. The state
validator also accepted unknown role keys, so a typo passed validation
and was ignored. The frontmatter generator now strips quoted spellings
of the key, which YAML's last-key-wins would otherwise honor over the
generated one.

Regression tests cover each case. The retry test drives the real Tier B
path and was verified to fail with the fix removed.

Quest/Co-Authored by
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: GPT-6 Astra <noreply@openai.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
Astra re-reviewed its own five fixes adversarially and found more:
snapshot migration still flattened PARTIAL effort maps to defaults, an
explicit `"effort": null` resolved as absent, unreadable config was
treated as missing, and rejecting an effort level truncated existing
artifacts because validation ran after artifact preparation. It also
added a bounded guard for finding 4 rather than converting native
dispatch to the runner, and fixed a pre-existing solo-mode migration
bug it hit while testing.

Reviewing that work turned up two over-corrections, both fixed here.

The native guard blocked any role the quest does not pin. Since the
generated frontmatter always names a level and an unpinned role never
can, that walled off every legacy quest on resume — the exact case the
previous two commits existed to protect. An unpinned role now passes;
only a genuine disagreement blocks.

The guard also rejected any agent file with nested YAML, so adding a
valid `hooks:` or `mcpServers:` block would have hard-blocked dispatch.
Nested blocks are now skipped, while a stray `effort:` at any depth
still fails closed as ambiguous.

Validating effort early also started raising out of run_claude_role,
replacing the JSON envelope orchestrators parse with a traceback. The
ordering fix is kept; the structured invocation_error is restored.

Quest/Co-Authored by
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: GPT-6 Astra <noreply@openai.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
@KjellKod
KjellKod marked this pull request as ready for review September 24, 2026 23:03
@KjellKod
KjellKod deployed to codex-ci-review September 24, 2026 23:03 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 24 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread .ai/schemas/allowlist.schema.json
Comment thread scripts/quest_sync_model_defaults.py Outdated
Comment thread scripts/quest_claude_runner.py
Comment thread scripts/quest_runtime/orchestration.py
Comment thread tests/unit/test_quest_effort.py Outdated
Comment thread .ai/quest.md Outdated
Comment thread scripts/quest_sync_model_defaults.py
Comment thread scripts/quest_runtime/orchestration.py
Comment thread scripts/quest_runtime/orchestration.py Outdated
Comment thread .skills/quest/delegation/workflow.md Outdated
Quest/Co-Authored by
Co-Authored-By: GPT-6 <noreply@openai.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
@KjellKod
KjellKod deployed to codex-ci-review September 24, 2026 23:20 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 10 files (changes from recent commits).

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread .ai/quest.md Outdated
Quest/Co-Authored by
Co-Authored-By: GPT-6 <noreply@openai.com>
in collaboration with KjellKod <kjell.hedstrom@gmail.com>
@KjellKod
KjellKod deployed to codex-ci-review September 24, 2026 23:29 — with GitHub Actions Active

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 1 file (changes from recent commits).

Requires human review: Wires per-role effort through dispatch, moves Claude roles to Opus 5.5, and changes review_mode/plan-iteration defaults. Human review required: native Task() effort is explicitly unverified; model/effort defaults are spend/quality tradeoffs.

Re-trigger cubic

@KjellKod KjellKod closed this Oct 1, 2026
@KjellKod KjellKod mentioned this pull request Oct 1, 2026
13 of 15 tasks

This branch was successfully deployed

1 active deployment
codex-ci-review — 4e4aa756 Deployed Sep 24, 2026 by KjellKod via codex-review #741
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant