Repository navigation
test: reconcile release regressions with typed planning owners - #6195
Conversation
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; runtime_reported; reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
精确 head 5e4b51e145e7ab96710f7532ac746d6b82b43d32;本结论批准已有回归的修复,不认证完整 1.4.0 发布。
动机
发布维护者和依赖回归验证判断是否可升级的开发者。主线已迁移任务规划规则,但旧 smoke 仍调用删除的 Python helper,PR 等待测试把缺失合并证据误当作仍在等待;有用的新指引也超过旧字符预算,阻碍发布判断。现有 smoke 直接经过当前类型化规划 owner;Python 等待断言与已接纳规则一致,预算按同负载测量调整,原结算、隐私与重放断言保留。不更改生产状态机、SQLite 默认、权限、调度或模型资格,不把这批回归修复称为完整发布通过。完整发布的全量测试、真实模型矩阵、发布物和个人指南仍有失败或未验证项,本 PR 不关闭这些缺口。
改动思路
先沿用已接纳的规则,再修正验证调用。四项 smoke 经 Python 事实编码调用 TypeScript 的统一 quota planning,保留原 lane 的数量、排序、身份和 scope 断言;不恢复已删除的 Python helper。PR 的 repo/number 只标识观察对象,没有合并事件并不能证明仍在等待。凭据 census 移除已经退出的 Python 站点,未来重新引入规则仍会被扫描为 offender。
预算按 docs/development/testing-and-quality.md#budget-failure-decisions 判断。v1.3.1 7445ae360 与当前相同稳定 fixture,writeback 从 1,831 增至 5,328 字符:完整 authoring 指引和 monitor 不可用时原身份结算说明仍有消费者价值。仅该阶段改为 6,000,剩余 672;spend 和 settled 继续 2,000。另一相同 36-Todo/12-run fixture 在 f96be230c 为 44,894,当前 45,249,差异是 cadence owner 的有效条件读回;具体 authoring allowance 改为 46,000,剩余 751。没有缩减 fixture、删除语义或调整任何冻结 benchmark/SQLite 晋升阈值。
具体改动
规范来源 docs/reference/todo-continuation-readback.md,基线 0df53a13f20765412b4c51026d1f28fc2f356d69:Todo continuation/readback、typed handoff retirement 和 budget-failure decisions 的相关要求均 implemented;Pending-target evidence:implemented(PR 缺失合并事实仍受监督);Budget decision:implemented(同负载测量及语义保留);Single typed owner:implemented(原 smoke 经过当前 batch)。完整发布资格 deferred,归原 release 维护边界。
quota-todo-summary-readmodel与todo-user-gate-readmodel从实际user_gateowner 导入原函数,其拒绝/识别断言保留。quota-cleared-blocker-successor-gate和todo-route-continuation-lanes经过project_quota_planning读取handoff_lanes/route_lanes,不增加生产包装;全源 fixture 和原 scope/order/count 断言保留。test_delivery_response.py将不具备 pending 证明的 PR case 从正例矩阵移至两种 PR 绑定的明确负例;monitor、capacity、注入时钟等正例与篡改负例保留。绑定字段也被独立断言。test_monitor_observation_admission.py只对 durable writeback 使用测量后的预算,原 primary/Turn 身份、不能新增 delivery、阶段、重放与一次性结算全部保留。test_cli_output_budget.py只更新具体 authoring allowance,原冷路径、内联字段和可操作指引检查保留,通用 lane/scale 预算未放松。test_public_safety_credential_caller_faces.py删除已经迁走的handoff_note.pyexemption,不改变 AST 探测器或产品隐私规则。
对主干的风险
八个既有验证文件 +59/-26,生产差异为零。主要风险是预算掩盖增长,因此保留旧失败、同负载测量、未变的其他阶段/scale ceiling 与完整语义/负例。111 个精确 head Python 用例、四项隔离 HOME 的公开 smoke、六个 TypeScript delivery-response 用例通过;包括真实隔离 SQLite 的回执与结算。Ruff、完整语义检查、8 路径公开边界、语法/diff 与风险 canary 11 项通过。首次 canary 因未安装 TypeScript npm 开发依赖失败,准备仓库声明的依赖后复跑通过;没有查询 CI。更广的发布测试和模型失败保留在发布边界,不能被本结论覆盖。
我的整体评价
修复了明确的验证失效,未扩大产品范围。原生产规则及用户 UI/Lark/CLI 操作不变,长期结算语义保留。Future-facing pass 已应用:让旧验证复用统一 typed owner,退出 stale exception;无需新抽象。没有阻塞发现;按仓库 test-only 政策,须发布本 head 评审并获得 merge readiness 后才可用已有维护者授权合并。
English verdict: APPROVE - 5e4b51e145e7ab96710f7532ac746d6b82b43d32. Existing validation adopts accepted typed owners and measured regression allowances without runtime/default changes. 111 focused Python tests, four isolated public smokes, six Node tests, semantic/public-boundary and 11 risk checks passed; full release qualification remains separate.
Release validation exposed stale callers after the typed Todo planning migration: four public smokes still imported retired helpers, the Python PR-wait fixture treated an unobserved repository/number as pending evidence, and the credential-rule census retained a retired Python decision site. This updates those existing regressions to the accepted owners; production behavior and new-Goal SQLite/hard_lease defaults are unchanged.
The stable auxiliary-resume fixture measures 1,831 → 5,328 JSON characters for writeback, so its regression allowance becomes 6,000 while spend/settled packets keep 2,000. The same 36-Todo/12-run authoring fixture measures 44,894 → 45,249 characters with the added cadence readback, so its fixed allowance becomes 46,000. Guidance, fixture populations, identity, privacy, replay and settlement assertions remain intact; no frozen benchmark or SQLite promotion threshold changes.
Validation: 111 focused Python tests and four public smokes passed, including real isolated SQLite settlement. Ruff, full semantic vocabulary/inventory, syntax/diff and changed-path public-boundary checks passed. Risk canary passed all 11 selected checks after preparing its npm development prerequisite; an initial run failed solely because that prerequisite was missing. The broader release qualification retains its failed and untested results and is not certified by this test repair.
Written by: model_agent (OpenAI Codex, GPT-6), under maintainer direction. Existing accepted Todo continuation/readback, typed handoff retirement and budget-decision contracts at base
0df53a13f20765412b4c51026d1f28fc2f356d69are the review frame. Future-facing pass: migrate validation callers to the shared typed batch and remove the stale census exemption; no compatibility wrapper or new framework.