Skip to content

fix(curation): draw measurements_latest from each check's latest run - #694

Merged
kstonekuan merged 3 commits into
Hebbian-Robotics:mainfrom
VARUN3WARE:fix/measurements-latest-follows-latest-check-run
Oct 6, 2026
Merged

kstonekuan merged 3 commits into
Hebbian-Robotics:mainfrom
VARUN3WARE:fix/measurements-latest-follows-latest-check-run

Conversation

@VARUN3WARE

Copy link
Copy Markdown
Contributor

Fixes #693

Summary

measurements_latest now takes measurements only from the latest run of each (episode_id, check_name), as observations_latest already does. A key that the latest run of its check did not record is withdrawn. The cases are a newer version omitting the key and a latest run that errored. The wide episodes view keeps the column and reads NULL, so curation stops selecting on the old value, and snapshot export stops shipping it.

On the #693 reproduction (a real staged ingest through hflow.App):

before after
measurements_latest [('blur_score', '1', 0.9), ('frames', '2', 100.0)] [('frames', '2', 100.0)]
SELECT blur_score, frames FROM episodes [(0.9, 100.0)] [(None, 100.0)]
WHERE blur_score > 0.8 (1,) (0,)
exported snapshot measurements.parquet includes v1 blur_score [('frames', '2', 100.0)], verifies ok

Why

The view ranked every (episode_id, key) on its own. A key the newest run did not write had no newer row to lose to, so the first run that ever wrote it kept winning forever. Meanwhile check_runs_latest, which coverage counts from, reported the newer version. Omitting a key is the documented way for a check to say "no value". So any version bump that stops a check emitting a bad value left that value in place on every episode already cataloged.

Changes, all in src/hflow/curation.py:

  • measurements_latest joins check_runs_latest on (episode_id, run_fingerprint, check_name, check_version). This is safe for every row: Catalog writes each measurement in the same loop as its check_runs row, with the same identifiers, artifact/ keys included. check_runs_latest moves above it and inherits the catalog: crash-mid-repair after winning episodes race can leave dependents permanently stale #51 ranking comment, since it ranks by the episode's recorded_at the same way.
  • The per-key rank is kept on top of that join. Without it, a key that moved from one check to another would return two rows, one from each check's latest run. With it, the view still keeps its documented one row per (episode_id, key), and the "one key, one owner" tie behavior is unchanged.
  • The wide view's columns still come from every completed run's keys, which is the same set as before. A query naming a withdrawn key therefore reads NULL instead of failing to bind. The case-collision guard also keeps covering every append, as docs/CATALOG.md promises: "conflicts across episodes, appends, or older catalogs cause curation and snapshot export to refuse".
  • The refresh path drops measurements_latest before check_runs_latest, since the first now depends on the second.

Behavior when the latest run errored (the question raised in #693): the errored run withdraws that check's earlier measurements. This matches observations_latest and the coverage rule ("an error on replay withdraws the earlier run's coverage"). It also removes a mismatch where coverage reported a step as not run while episodes still showed its value. If you would rather keep values through a transient error, it is a small change to the same view: rank only over runs with status passed, failed or measured. test_an_errored_latest_run_withdraws_that_checks_measurements pins the current choice.

No stored data changes. The fix only changes what the views read, so existing catalogs pick it up on the next open.

Validation

uv run ruff check                  # All checks passed
uv run ruff format --check         # 323 files already formatted
uv run ty check                    # All checks passed
uv run pytest -q -n auto           # 2518 passed, 8 skipped (includes packages/hflow-server/tests)

New tests in tests/test_catalog_curation.py:

  • test_a_key_a_newer_check_version_omits_is_withdrawn: the omitted key leaves measurements_latest; the wide view reads NULL; a check absent from the newer run keeps its value; curation no longer selects on it.
  • test_an_errored_latest_run_withdraws_that_checks_measurements
  • test_a_key_moved_to_another_check_keeps_one_latest_row

Mutation checks, each run against the new tests:

  • The original curation.py: the omitted-key and errored-run tests fail. The moved-key test passes, since the old per-key rank already handled that case.
  • Plain join without the per-key rank: the moved-key test fails with two rows, old_check and new_check.
  • Wide-view columns built from measurements_latest: queries naming a withdrawn key fail with a binder error, and the existing cross-append case-collision test stops refusing.

I also ran the #693 reproduction end to end, plus snapshot export and verify, with the before and after results in the table above.

Checklist

  • I added or updated outcome-focused tests for changed business logic.
  • I updated documentation for changed behavior, flags, formats, or requirements.
  • I ran uv run ruff check --fix, uv run ruff format, and uv run ty check.
  • I ran the relevant pytest suite.
  • I did not add recordings, generated media, credentials, private URLs, or runtime artifacts.
  • I preserved stored-data compatibility or documented an explicit version change.

measurements_latest ranked every (episode_id, key) on its own, so a key the latest run of its check no longer recorded kept the older run's value forever. The wide episodes view, curation, and snapshot export then served a withdrawn value while check_runs_latest reported the newer version.

Join measurements through check_runs_latest, as observations_latest already does, and keep the per-key rank so a key that moved between checks still yields one row. The wide view's columns stay enumerated from every completed run's keys: a withdrawn key reads NULL instead of failing to bind, and the case-collision guard still covers every append.

Refs Hebbian-Robotics#693
… moved keys

Covers a key a newer check version omits (withdrawn from measurements_latest and the wide view, untouched checks keep theirs, curation no longer selects on it), an errored latest run withdrawing its check's measurements, and a key moved to another check still yielding one latest row.

Refs Hebbian-Robotics#693
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Critical risk] Changes how measurement data is selected from check runs.

The PR appears safe to merge; the review found no actionable regression.

What we checked:

  • Measurements lose their check-run link: The catalog writer inserts a check-run row and its measurements together, using the same run identifiers.
  • Catalog refresh drops views in the wrong order: Refresh drops measurements_latest before check_runs_latest, which it now uses.
Summary

The catalog now shows measurements only from each check’s latest run, so omitted keys and errored runs no longer leave old values available to curation or snapshot export. The wide episodes view keeps withdrawn keys as columns with NULL values.

  • Measurements come only from each check’s latest run.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[check_runs and episodes_raw] --> B[check_runs_latest]
  C[measurements] --> D[measurements_latest]
  B --> D
  D --> E[wide episodes view]
  D --> F[snapshot measurements]
Loading

Reviews (1) · Last reviewed commit: "docs(catalog): state that measurements_l..."

@kstonekuan kstonekuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. This matches observations_latest and the withdrawal semantics from #676, and the full suite passes merged onto main. Merging once CI finishes.

@kstonekuan
kstonekuan merged commit a7952c3 into Hebbian-Robotics:main Oct 6, 2026
6 of 7 checks passed
@kstonekuan kstonekuan mentioned this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants