|
1 | 1 | # Changelog |
2 | 2 |
|
| 3 | +## v0.18.2 — 2026-08-25 |
| 4 | + |
| 5 | +Budgets that were spent by failure. |
| 6 | + |
| 7 | +v0.18.1 left one measured defect unsolved: `fix-a-bug` hit its 20-minute ceiling |
| 8 | +in one run of three with `failed_tasks == 0` — nothing wrong except that nobody |
| 9 | +asked whether the work was done. Chasing it found the cause, and the cause |
| 10 | +turned out to be a **pattern** rather than a bug. Six mechanisms in this |
| 11 | +codebase bounded themselves with a fixed count, charged that count for negative |
| 12 | +or failed attempts, and so switched themselves off precisely on the runs that |
| 13 | +needed them. Every one of them failed silently. |
| 14 | + |
| 15 | +**Measured, three runs per scenario, Qwen3.5-9B on oMLX, same fixtures and the |
| 16 | +same 20-minute ceiling as the v0.18.1 measurement:** |
| 17 | + |
| 18 | +| | v0.18.1 | v0.18.2 | |
| 19 | +|---|---|---| |
| 20 | +| `implement-from-tests` median prompt tokens | 123,652 | **75,535 (−39%)** | |
| 21 | +| `implement-from-tests` pass rate | 3 of 3 | 3 of 3 | |
| 22 | +| `fix-a-bug` median prompt tokens | ~178,000 | ~182,000 | |
| 23 | +| `fix-a-bug` pass rate | 2 of 3 | 2 of 3 | |
| 24 | +| runs terminating within budget | 5 of 6 | 5 of 6 | |
| 25 | + |
| 26 | +**Read that table honestly: one scenario improved substantially and the other |
| 27 | +did not move.** The probe-starvation defect was real, is fixed, and is confirmed |
| 28 | +firing in a live run — the run metrics record |
| 29 | +`"gates":[{"name":"qa_gate","passed":true},{"name":"objective_met_early","passed":true}]`, |
| 30 | +which is the between-waves probe ending a run the moment the objective went |
| 31 | +green, in 6 LLM calls and 375s. But it was **not** what `fix-a-bug` was failing |
| 32 | +on, and that scenario's 1-in-3 ceiling miss is unchanged. |
| 33 | + |
| 34 | +What that remaining miss actually is, measured: the same 9B needs 8 tool calls |
| 35 | +and 69k tokens for this three-line boundary fix on a good run and **32 tool |
| 36 | +calls and 192k tokens** on a bad one. In the run that hits the ceiling the |
| 37 | +harness still returns a CORRECT result — `engine_success=true`, the fixture |
| 38 | +tests green, the protected test file byte-identical, `failed_tasks=0`. The only |
| 39 | +failing assertion is the suite's own wall-clock budget. So this is model |
| 40 | +variance, not a harness that cannot tell it is done. |
| 41 | + |
| 42 | +The identified next step, with evidence behind it and deliberately NOT taken |
| 43 | +here: the probe fires only between waves, so when a worker fixes the bug at tool |
| 44 | +call 10 of 32, nothing notices until the task ends. That worker runs the |
| 45 | +objective command itself through `ws_shell` during those calls, and the harness |
| 46 | +already sees the output — a green result there is free evidence the probe could |
| 47 | +harvest without spending anything. It is a real design and an unproven one, and |
| 48 | +this release does not ship unproven changes. |
| 49 | + |
| 50 | +Wall-clock is deliberately absent from the table above. Across these runs the |
| 51 | +same suite showed FEWER tokens taking MORE seconds (75k/792s here against |
| 52 | +124k/608s for v0.18.1), i.e. throughput degraded over an hour of continuous |
| 53 | +local GPU load. Wall times are not comparable between sessions on this hardware; |
| 54 | +token counts do not depend on machine speed, so they are what is reported. |
| 55 | + |
| 56 | +### Fixed |
| 57 | + |
| 58 | +- **The objective probe no longer goes blind three waves into a run.** |
| 59 | + `maxObjectiveProbes` was 2, and the budget was charged for RED answers. |
| 60 | + Between-waves probing spends it on the earliest waves — the two moments in a |
| 61 | + run when "not yet" is most certain and least worth paying to learn — so from |
| 62 | + wave three nothing ever asked again. If the implementation landed at wave |
| 63 | + five, the run burned its whole ceiling with the work already complete. The |
| 64 | + budget is also **per run, not per board**: a run drives `RunBoard` once per |
| 65 | + corrective round, so two probes was two for the entire run, and the post-drain |
| 66 | + probe that corrective boards deliberately rely on (they run with early-stop |
| 67 | + off) was starved by the same exhaustion. |
| 68 | + A count also cannot be right for two projects at once — two probes is miserly |
| 69 | + for a 200ms unit suite and profligate for a 6-minute integration suite. The |
| 70 | + bound is now economic: the probe's cost is **measured** (`SmokeResult.Duration`, |
| 71 | + timed where the command actually runs) and asking continues while |
| 72 | + `runway >= 6 × cost`. A cheap gate is asked on every wave that wrote |
| 73 | + something; an expensive one only while there is runway for the answer to pay |
| 74 | + for itself. Neither number is guessed per project. |
| 75 | +- **A slow endpoint no longer loses structured decoding for seven days.** |
| 76 | + One deadline covers up to six sequential capability probes. A cold local model |
| 77 | + — the exact case the budget was widened for — could spend most of it loading |
| 78 | + weights for the first, after which every later probe died on the shared |
| 79 | + deadline. `attempt` returns false for a transport error exactly as for a 400, |
| 80 | + so the negotiation could not tell "the server refused this field" from "we |
| 81 | + never got to ask", and stamped `Source: "probe"` on a wholesale-false record. |
| 82 | + `capCache` then persisted and honored it for `CapabilityTTL` — seven days in |
| 83 | + which every structured role on that endpoint silently degraded to prompt-only |
| 84 | + + repair, with no path back, because nothing re-probes a record that is still |
| 85 | + fresh. A cut-short negotiation now falls back to the family preset and leaves |
| 86 | + `Probed` zero, which keeps it out of the on-disk cache. Preferring the preset |
| 87 | + over all-false is the recoverable direction: an over-claimed mechanism costs |
| 88 | + one 400 and is then recorded by `demoteCapability`; an under-claimed one has |
| 89 | + nothing that can ever notice it. |
| 90 | +- **Retrieval no longer goes blind when the corpus is most on-topic.** |
| 91 | + The relative noise floor is documented as "the corpus's own median", but it |
| 92 | + was measured over the value `Search` returns — already truncated to `TopK`. |
| 93 | + With the default `TopK=5` the "median" was the **third-best hit**, so the |
| 94 | + threshold became third-best + `NoiseMargin`: arithmetically at most two hits |
| 95 | + could ever survive, and a tightly clustered top five — a set of uniformly |
| 96 | + strong matches — cleared nothing at all and `RetrieveForQuery` returned `""` |
| 97 | + with no error and no warning. It also made `MinChunksForNoiseFloor` |
| 98 | + unreachable for its stated purpose, since it compared against `len(top-k)` and |
| 99 | + never the corpus size. `Retriever.SearchAll` now supplies the whole scored |
| 100 | + distribution for the floor, and the top-k truncation happens after. |
| 101 | +- **A cancelled QA gate no longer reports a red verdict.** `runQAGate` returns |
| 102 | + "the gate failed", and every cancellation path returned **true**. |
| 103 | + `finalizeAfterExecute` reads that as `QAFailed`, sets `TesterRejected`, and |
| 104 | + feeds the board a synthesized tester verdict |
| 105 | + (`{"passed":false,...,"qa_gate red"}`) through `applyTesterFeedback` — a |
| 106 | + planner call. So Ctrl-C, or a scenario budget expiring, bought an extra LLM |
| 107 | + round-trip during shutdown and annotated a done task with "QA gate still |
| 108 | + failing" about a gate never allowed to finish. A cancelled gate now records |
| 109 | + nothing: no verdict, no annotation. |
| 110 | +- **The QA gate stops re-running an unchanged tree.** `qaDiagnoseAndFix` |
| 111 | + discarded both role outputs and never reported whether anything was written, |
| 112 | + so a fix pass that produced nothing (budget exhausted, a refusal, a prose-only |
| 113 | + answer) was followed by a byte-identical command run against a byte-identical |
| 114 | + tree, at full price, for every remaining round. The objective probe has |
| 115 | + refused exactly this since it was written; the gate never learned it. Measured |
| 116 | + on a stalled gate: 4 command runs down to 1. Gate rounds also now price the |
| 117 | + command for the probe budget, since they run the same one. |
| 118 | +- **A cold start no longer pins concurrency for a month.** The calibration probe |
| 119 | + runs inside a fixed wall-clock budget in which every unit of work is a model |
| 120 | + call — the thing being measured. On a slow server the warm-up and solo |
| 121 | + baseline can exhaust it before any concurrency level is measured, leaving only |
| 122 | + the synthetic single-level entry, from which `SelectKnee` returns 1. |
| 123 | + `Profile.Current` checked ID, `MaxParallel`, version and age but **not |
| 124 | + `Partial`**, so that degenerate verdict was served from cache for |
| 125 | + `DefaultTTL`; `Apply` only checks `MaxParallel > 0`, and the "partial" marker |
| 126 | + appears solely in `Summary()`, which the auto path never prints. So the |
| 127 | + slowest models — the ones with the most to gain from a real measurement — were |
| 128 | + silently capped at `max_parallel=1` for thirty days. Partial profiles now |
| 129 | + expire after `PartialTTL` (1 hour): a cold start is transient, so the retry |
| 130 | + should be too. |
| 131 | +- **The thrash detector no longer blinds itself on thrashing runs.** The |
| 132 | + signature map **froze** at `MaxSignatures` — once full, a signature it had not |
| 133 | + already seen was never admitted, so its later repeats were never counted. The |
| 134 | + asymmetry is what makes it bite: the classic small-model edit failure is a |
| 135 | + near-miss `ws_edit` retried with a slightly different `old_str` each time, and |
| 136 | + every one of those is a *distinct* signature, so a thrashing run fills the map |
| 137 | + faster than a healthy one and goes blind sooner. Measured: 5 identical calls |
| 138 | + after 256 distinct ones counted **0** repeats instead of 4. |
| 139 | + `RunReport.RedundantCalls` and the redundant-call-rate KPI under-reported with |
| 140 | + nothing marking the count as capped, and evolve learned from the truncated |
| 141 | + signal. It now evicts oldest-first, like every other bound in that file. |
| 142 | +- **A reviewer that timed out is no longer recorded as one that judged the work |
| 143 | + worthless.** The synthesized error attempt carries `NoVerdict`, and |
| 144 | + `Attempt.Score` documents that a 0 under an `error` verdict is an absence, not |
| 145 | + a judgement. |
| 146 | +- **`autoresearch` no longer claims a surface was exhausted when it was not.** |
| 147 | + `StopExhausted` is reached on `ErrNoProposal` from the *deterministic* |
| 148 | + proposer, which cannot touch a text knob at all — yet its sentence read "every |
| 149 | + value of every knob was tried", the one message in that package built to be |
| 150 | + trusted, while every sibling carefully says when the surface was *not* |
| 151 | + exhausted. |
| 152 | + |
| 153 | +### Known, reported rather than changed |
| 154 | + |
| 155 | +Three further findings are real and measured but are **behaviour changes whose |
| 156 | +benefit cannot be proven without a controlled study**, so they are documented |
| 157 | +instead of shipped unmeasured: |
| 158 | + |
| 159 | +- `reviewerStrictDelay` (20ms) means the default `max_parallel=4` issues **two |
| 160 | + reviewer LLM requests per review**, and `strictOut` is used only when the |
| 161 | + primary reply is empty — so the second is discarded almost always. The code is |
| 162 | + honest about the cost (`noteExtraRequests` reports it), but on a local server |
| 163 | + that runs inference serially this roughly doubles review latency. A delay |
| 164 | + derived from measured reviewer p50 would keep the insurance and drop the cost. |
| 165 | +- `readBudgetLines` sizes `ws_read` from `MaxContextKB`, the legacy prompt-byte |
| 166 | + budget, rather than the model's real `ContextLimit` — the exact conflation |
| 167 | + `compact.WindowTokensFromKB` is already marked Deprecated for. On typical |
| 168 | + source that caps a read around 80–120 lines whatever the model's real window |
| 169 | + is. Fixing it would raise read sizes several-fold on a large-context model, |
| 170 | + which needs measuring before it ships. |
| 171 | +- `agents.factory` charges loop-guard interventions against the ReAct iteration |
| 172 | + budget (both escalation paths return a successful tool result), and a model |
| 173 | + profile's `max_turns` can only *lower* the default 8, never raise it. |
| 174 | + |
3 | 175 | ## v0.18.1 — 2026-08-25 |
4 | 176 |
|
5 | 177 | The harness stops when the work is done. |
|
0 commit comments