Repository navigation
fix(curation): draw measurements_latest from each check's latest run - #694
Merged
kstonekuan merged 3 commits intoOct 6, 2026
Conversation
measurements_latest ranked every (episode_id, key) on its own, so a key the latest run of its check no longer recorded kept the older run's value forever. The wide episodes view, curation, and snapshot export then served a withdrawn value while check_runs_latest reported the newer version. Join measurements through check_runs_latest, as observations_latest already does, and keep the per-key rank so a key that moved between checks still yields one row. The wide view's columns stay enumerated from every completed run's keys: a withdrawn key reads NULL instead of failing to bind, and the case-collision guard still covers every append. Refs Hebbian-Robotics#693
… moved keys Covers a key a newer check version omits (withdrawn from measurements_latest and the wide view, untouched checks keep theirs, curation no longer selects on it), an errored latest run withdrawing its check's measurements, and a key moved to another check still yielding one latest row. Refs Hebbian-Robotics#693
Contributor
|
kstonekuan
approved these changes
Oct 6, 2026
kstonekuan
left a comment
Contributor
There was a problem hiding this comment.
LGTM. This matches observations_latest and the withdrawal semantics from #676, and the full suite passes merged onto main. Merging once CI finishes.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #693
Summary
measurements_latestnow takes measurements only from the latest run of each(episode_id, check_name), asobservations_latestalready does. A key that the latest run of its check did not record is withdrawn. The cases are a newer version omitting the key and a latest run that errored. The wideepisodesview keeps the column and readsNULL, so curation stops selecting on the old value, and snapshot export stops shipping it.On the #693 reproduction (a real staged ingest through
hflow.App):measurements_latest[('blur_score', '1', 0.9), ('frames', '2', 100.0)][('frames', '2', 100.0)]SELECT blur_score, frames FROM episodes[(0.9, 100.0)][(None, 100.0)]WHERE blur_score > 0.8(1,)(0,)measurements.parquetblur_score[('frames', '2', 100.0)], verifiesokWhy
The view ranked every
(episode_id, key)on its own. A key the newest run did not write had no newer row to lose to, so the first run that ever wrote it kept winning forever. Meanwhilecheck_runs_latest, which coverage counts from, reported the newer version. Omitting a key is the documented way for a check to say "no value". So any version bump that stops a check emitting a bad value left that value in place on every episode already cataloged.Changes, all in
src/hflow/curation.py:measurements_latestjoinscheck_runs_lateston(episode_id, run_fingerprint, check_name, check_version). This is safe for every row:Catalogwrites each measurement in the same loop as itscheck_runsrow, with the same identifiers,artifact/keys included.check_runs_latestmoves above it and inherits the catalog: crash-mid-repair after winning episodes race can leave dependents permanently stale #51 ranking comment, since it ranks by the episode'srecorded_atthe same way.(episode_id, key), and the "one key, one owner" tie behavior is unchanged.NULLinstead of failing to bind. The case-collision guard also keeps covering every append, asdocs/CATALOG.mdpromises: "conflicts across episodes, appends, or older catalogs cause curation and snapshot export to refuse".measurements_latestbeforecheck_runs_latest, since the first now depends on the second.Behavior when the latest run errored (the question raised in #693): the errored run withdraws that check's earlier measurements. This matches
observations_latestand the coverage rule ("an error on replay withdraws the earlier run's coverage"). It also removes a mismatch where coverage reported a step as not run whileepisodesstill showed its value. If you would rather keep values through a transient error, it is a small change to the same view: rank only over runs with statuspassed,failedormeasured.test_an_errored_latest_run_withdraws_that_checks_measurementspins the current choice.No stored data changes. The fix only changes what the views read, so existing catalogs pick it up on the next open.
Validation
New tests in
tests/test_catalog_curation.py:test_a_key_a_newer_check_version_omits_is_withdrawn: the omitted key leavesmeasurements_latest; the wide view readsNULL; a check absent from the newer run keeps its value; curation no longer selects on it.test_an_errored_latest_run_withdraws_that_checks_measurementstest_a_key_moved_to_another_check_keeps_one_latest_rowMutation checks, each run against the new tests:
curation.py: the omitted-key and errored-run tests fail. The moved-key test passes, since the old per-key rank already handled that case.old_checkandnew_check.measurements_latest: queries naming a withdrawn key fail with a binder error, and the existing cross-append case-collision test stops refusing.I also ran the #693 reproduction end to end, plus snapshot export and verify, with the before and after results in the table above.
Checklist
uv run ruff check --fix,uv run ruff format, anduv run ty check.