Skip to content

test: reconcile release regressions with typed planning owners - #6195

Merged
loopx-agent merged 1 commit into
mainfrom
codex/release-140-current-contract-tests
Oct 11, 2026
Merged

loopx-agent merged 1 commit into
mainfrom
codex/release-140-current-contract-tests

Conversation

@loopx-agent

@loopx-agent loopx-agent commented Oct 11, 2026 •

Copy link
Copy Markdown
Collaborator

Release validation exposed stale callers after the typed Todo planning migration: four public smokes still imported retired helpers, the Python PR-wait fixture treated an unobserved repository/number as pending evidence, and the credential-rule census retained a retired Python decision site. This updates those existing regressions to the accepted owners; production behavior and new-Goal SQLite/hard_lease defaults are unchanged.

The stable auxiliary-resume fixture measures 1,831 → 5,328 JSON characters for writeback, so its regression allowance becomes 6,000 while spend/settled packets keep 2,000. The same 36-Todo/12-run authoring fixture measures 44,894 → 45,249 characters with the added cadence readback, so its fixed allowance becomes 46,000. Guidance, fixture populations, identity, privacy, replay and settlement assertions remain intact; no frozen benchmark or SQLite promotion threshold changes.

Validation: 111 focused Python tests and four public smokes passed, including real isolated SQLite settlement. Ruff, full semantic vocabulary/inventory, syntax/diff and changed-path public-boundary checks passed. Risk canary passed all 11 selected checks after preparing its npm development prerequisite; an initial run failed solely because that prerequisite was missing. The broader release qualification retains its failed and untested results and is not certified by this test repair.

Written by: model_agent (OpenAI Codex, GPT-6), under maintainer direction. Existing accepted Todo continuation/readback, typed handoff retirement and budget-decision contracts at base 0df53a13f20765412b4c51026d1f28fc2f356d69 are the review frame. Future-facing pass: migrate validation callers to the shared typed batch and remove the stale census exemption; no compatibility wrapper or new framework.

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; runtime_reported; reasoning_effort=xhigh

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

精确 head 5e4b51e145e7ab96710f7532ac746d6b82b43d32;本结论批准已有回归的修复,不认证完整 1.4.0 发布。

动机

发布维护者和依赖回归验证判断是否可升级的开发者。主线已迁移任务规划规则,但旧 smoke 仍调用删除的 Python helper,PR 等待测试把缺失合并证据误当作仍在等待;有用的新指引也超过旧字符预算,阻碍发布判断。现有 smoke 直接经过当前类型化规划 owner;Python 等待断言与已接纳规则一致,预算按同负载测量调整,原结算、隐私与重放断言保留。不更改生产状态机、SQLite 默认、权限、调度或模型资格,不把这批回归修复称为完整发布通过。完整发布的全量测试、真实模型矩阵、发布物和个人指南仍有失败或未验证项,本 PR 不关闭这些缺口。

改动思路

先沿用已接纳的规则,再修正验证调用。四项 smoke 经 Python 事实编码调用 TypeScript 的统一 quota planning,保留原 lane 的数量、排序、身份和 scope 断言;不恢复已删除的 Python helper。PR 的 repo/number 只标识观察对象,没有合并事件并不能证明仍在等待。凭据 census 移除已经退出的 Python 站点,未来重新引入规则仍会被扫描为 offender。

预算按 docs/development/testing-and-quality.md#budget-failure-decisions 判断。v1.3.1 7445ae360 与当前相同稳定 fixture,writeback 从 1,831 增至 5,328 字符:完整 authoring 指引和 monitor 不可用时原身份结算说明仍有消费者价值。仅该阶段改为 6,000,剩余 672;spend 和 settled 继续 2,000。另一相同 36-Todo/12-run fixture 在 f96be230c 为 44,894,当前 45,249,差异是 cadence owner 的有效条件读回;具体 authoring allowance 改为 46,000,剩余 751。没有缩减 fixture、删除语义或调整任何冻结 benchmark/SQLite 晋升阈值。

具体改动

规范来源 docs/reference/todo-continuation-readback.md,基线 0df53a13f20765412b4c51026d1f28fc2f356d69:Todo continuation/readback、typed handoff retirement 和 budget-failure decisions 的相关要求均 implemented;Pending-target evidence:implemented(PR 缺失合并事实仍受监督);Budget decision:implemented(同负载测量及语义保留);Single typed owner:implemented(原 smoke 经过当前 batch)。完整发布资格 deferred,归原 release 维护边界。

  • quota-todo-summary-readmodel 与 todo-user-gate-readmodel 从实际 user_gate owner 导入原函数,其拒绝/识别断言保留。
  • quota-cleared-blocker-successor-gate 和 todo-route-continuation-lanes 经过 project_quota_planning 读取 handoff_lanes/route_lanes,不增加生产包装;全源 fixture 和原 scope/order/count 断言保留。
  • test_delivery_response.py 将不具备 pending 证明的 PR case 从正例矩阵移至两种 PR 绑定的明确负例;monitor、capacity、注入时钟等正例与篡改负例保留。绑定字段也被独立断言。
  • test_monitor_observation_admission.py 只对 durable writeback 使用测量后的预算,原 primary/Turn 身份、不能新增 delivery、阶段、重放与一次性结算全部保留。
  • test_cli_output_budget.py 只更新具体 authoring allowance,原冷路径、内联字段和可操作指引检查保留,通用 lane/scale 预算未放松。
  • test_public_safety_credential_caller_faces.py 删除已经迁走的 handoff_note.py exemption,不改变 AST 探测器或产品隐私规则。

对主干的风险

八个既有验证文件 +59/-26,生产差异为零。主要风险是预算掩盖增长,因此保留旧失败、同负载测量、未变的其他阶段/scale ceiling 与完整语义/负例。111 个精确 head Python 用例、四项隔离 HOME 的公开 smoke、六个 TypeScript delivery-response 用例通过;包括真实隔离 SQLite 的回执与结算。Ruff、完整语义检查、8 路径公开边界、语法/diff 与风险 canary 11 项通过。首次 canary 因未安装 TypeScript npm 开发依赖失败,准备仓库声明的依赖后复跑通过;没有查询 CI。更广的发布测试和模型失败保留在发布边界,不能被本结论覆盖。

我的整体评价

修复了明确的验证失效,未扩大产品范围。原生产规则及用户 UI/Lark/CLI 操作不变,长期结算语义保留。Future-facing pass 已应用:让旧验证复用统一 typed owner,退出 stale exception;无需新抽象。没有阻塞发现;按仓库 test-only 政策,须发布本 head 评审并获得 merge readiness 后才可用已有维护者授权合并。

English verdict: APPROVE - 5e4b51e145e7ab96710f7532ac746d6b82b43d32. Existing validation adopts accepted typed owners and measured regression allowances without runtime/default changes. 111 focused Python tests, four isolated public smokes, six Node tests, semantic/public-boundary and 11 risk checks passed; full release qualification remains separate.

@loopx-agent
loopx-agent merged commit 8d5a6fc into main Oct 11, 2026
1 of 4 checks passed
@loopx-agent
loopx-agent deleted the codex/release-140-current-contract-tests branch October 11, 2026 06:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant