Skip to content

Add the hubverse model output interface - #14

Merged
nickreich merged 2 commits into
mainfrom
nr/hubverse-layer/3
Sep 2, 2026
Merged

nickreich merged 2 commits into
mainfrom
nr/hubverse-layer/3

Conversation

@nickreich

@nickreich nickreich commented Sep 2, 2026 •

Copy link
Copy Markdown
Member

The reason for the rewrite: hubverse model output goes in, allocation scores
come out.

alloscore_model_out() takes a model_out_tbl of quantile forecasts and the
matching oracle_output and scores them, replacing roughly thirty lines of
per-analysis boilerplate — nest ps/qs, deframe,
add_pdqr_funs(dist = "distfromq"), join truth — that was reproduced in every
downstream script. See utility-eval-papers/R/run-alloscore.R for the version
this displaces.

scores <- alloscore_model_out(
  model_out_tbl = dplyr::filter(hubExamples::forecast_outputs, output_type == "quantile"),
  oracle_output = hubExamples::forecast_oracle_output,
  K = c(500, 1000, 2000),
  target_cols = "location",
  by = c("model_id", "K")
)

The one genuinely new design decision

An allocation problem is a set of targets sharing one budget, and hubverse has
no single column for that. So target_cols names the task ID columns whose
combinations enumerate those targets, and the allocation unit is derived as
model_id plus every remaining task ID column. One problem is solved per
combination.

That mirrors how a hubverse compound_taskid_set works: the columns you name
vary within a group, the rest are held constant. With
target_cols = "location" on the hubExamples forecast data you get 24
allocation problems (3 models x 2 reference dates x 4 horizons), each pooling 2
locations.

This is the part most worth review attention — it is an interpretation of the
hubverse data model, not a mechanical port.

Functions

  • as_alloscore_df() — exported, because it is where a mis-specified join or
    target set becomes visible. Returns one row per allocation unit with a
    forecasts list column holding each target's quantiles, its distfromq-built
    cdf and quantile function, and the observed outcome.
  • allocate_model_out() — stops after the allocation, for when the question is
    what a forecast implies you should do rather than how good it was.
  • alloscore_model_out() — allocates and scores, with summarize/by in the
    shape hubEvals::score_model_out() uses.

Validation

Follows hubEvals: get_task_id_cols() and validate_model_oracle_out()
mirror its helpers, and a single supported output_type is enforced up front.
Only quantile is accepted; sample, cdf, pmf, mean and median error
with an explanation rather than silently doing something wrong.

Rejected inputs covered by tests: several output types at once, a non-quantile
output type, non-numeric quantile levels, oracle_output with unexpected
columns or missing coverage, duplicated forecasts, and target_cols naming
something that is not a task ID.

Tests

418 total, up from 306, in two commits.

The interface adds a test worth looking at,
"alloscore_model_out matches the core API on a single unit": it routes one
allocation problem through both the hubverse entry point and the core
alloscore() and asserts the scores agree, which is what makes this a refactor
of the interface rather than a second implementation.

The equivalence fixtures then prove both agree with the original package.
The original has no test suite, so there is no pre-existing ground truth to
port: data-raw/make_legacy_reference.R runs the original code once over a
fixed set of fixtures and freezes the results under
tests/testthat/testdata/, and test-legacy-equivalence.R checks this package
reproduces them. The original therefore stays out of this package's
dependencies while the numerics stay pinned, so any future change in behaviour
has to be deliberate.

legacy_hub_quantile_scores.csv is the one that matters here: hubverse example
data pushed through the original's hand-rolled pipeline, which
alloscore_model_out() must reproduce. It ships alongside the function rather
than in a later PR, so the interface does not land unproven.

legacy_reference_meta.json records the original's git SHA (0477ae9) and the
R, distfromq and hubExamples versions used. Tolerances are 1e-6 on
allocations and 1e-8 on scores. A mutation check confirms the suite is not
vacuous: widening eps_K from 1% to 5% fails these tests and the ZXH tests,
and restoring it passes all 54.

Verified locally

R CMD check: Status: OK. 418 tests, 0 failures, 0 warnings, 0 skips.
lintr: 0 lints. air format . --check: clean.

nickreich and others added 2 commits September 2, 2026 16:55
The reason for the rewrite: `alloscore_model_out()` takes a hubverse
`model_out_tbl` and `oracle_output` and scores them, replacing roughly thirty
lines of per-analysis boilerplate (nest `ps`/`qs`, `deframe`,
`add_pdqr_funs(dist = "distfromq")`, join truth) reproduced in every downstream
script — see `utility-eval-papers/R/run-alloscore.R`.

The one genuinely new design decision is grouping. An allocation problem is a
*set* of targets sharing one budget, which hubverse has no single column for, so
`target_cols` names the task ID columns whose combinations enumerate those
targets and the allocation unit is derived as `model_id` plus everything else.
That mirrors `compound_taskid_set`: the columns you name vary within a group,
the rest are held constant.

`as_alloscore_df()` is exported as well, since it is where a mis-specified join
or target set shows up.

Validation follows `hubEvals`: `get_task_id_cols()` and
`validate_model_oracle_out()` mirror its helpers, and a single supported
`output_type` is enforced up front. `test-model_out.R` includes a test that one
allocation problem routed through the hubverse API and through the core API
gives the same score.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The original package has no test suite, so there is no pre-existing ground
truth to port. data-raw/make_legacy_reference.R runs the original code once
over a fixed set of fixtures and freezes the results in
tests/testthat/testdata/; test-legacy-equivalence.R checks that this package
reproduces them. That keeps the original out of this package's dependencies
while still pinning the numerics, so any future change in behaviour has to be
deliberate.

The fixture that matters most is legacy_hub_quantile_scores.csv: hubverse
example data pushed through the original package's hand-rolled pipeline, which
alloscore_model_out() must reproduce. It ships in the same pull request as that
function so the interface does not land unproven.

legacy_reference_meta.json records the original package's git SHA (0477ae9) and
the R, distfromq and hubExamples versions everything was generated under.

Tolerances are 1e-6 on allocations and 1e-8 on scores, loose enough to absorb
uniroot noise and far tighter than any behavioural difference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@nickreich
nickreich merged commit c2d9de7 into main Sep 2, 2026
9 checks passed
@nickreich
nickreich deleted the nr/hubverse-layer/3 branch September 2, 2026 21:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant