Add the hubverse model output interface - #14
Merged
Merged
Conversation
The reason for the rewrite: `alloscore_model_out()` takes a hubverse `model_out_tbl` and `oracle_output` and scores them, replacing roughly thirty lines of per-analysis boilerplate (nest `ps`/`qs`, `deframe`, `add_pdqr_funs(dist = "distfromq")`, join truth) reproduced in every downstream script — see `utility-eval-papers/R/run-alloscore.R`. The one genuinely new design decision is grouping. An allocation problem is a *set* of targets sharing one budget, which hubverse has no single column for, so `target_cols` names the task ID columns whose combinations enumerate those targets and the allocation unit is derived as `model_id` plus everything else. That mirrors `compound_taskid_set`: the columns you name vary within a group, the rest are held constant. `as_alloscore_df()` is exported as well, since it is where a mis-specified join or target set shows up. Validation follows `hubEvals`: `get_task_id_cols()` and `validate_model_oracle_out()` mirror its helpers, and a single supported `output_type` is enforced up front. `test-model_out.R` includes a test that one allocation problem routed through the hubverse API and through the core API gives the same score. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The original package has no test suite, so there is no pre-existing ground truth to port. data-raw/make_legacy_reference.R runs the original code once over a fixed set of fixtures and freezes the results in tests/testthat/testdata/; test-legacy-equivalence.R checks that this package reproduces them. That keeps the original out of this package's dependencies while still pinning the numerics, so any future change in behaviour has to be deliberate. The fixture that matters most is legacy_hub_quantile_scores.csv: hubverse example data pushed through the original package's hand-rolled pipeline, which alloscore_model_out() must reproduce. It ships in the same pull request as that function so the interface does not land unproven. legacy_reference_meta.json records the original package's git SHA (0477ae9) and the R, distfromq and hubExamples versions everything was generated under. Tolerances are 1e-6 on allocations and 1e-8 on scores, loose enough to absorb uniroot noise and far tighter than any behavioural difference. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The reason for the rewrite: hubverse model output goes in, allocation scores
come out.
alloscore_model_out()takes amodel_out_tblof quantile forecasts and thematching
oracle_outputand scores them, replacing roughly thirty lines ofper-analysis boilerplate — nest
ps/qs,deframe,add_pdqr_funs(dist = "distfromq"), join truth — that was reproduced in everydownstream script. See
utility-eval-papers/R/run-alloscore.Rfor the versionthis displaces.
The one genuinely new design decision
An allocation problem is a set of targets sharing one budget, and hubverse has
no single column for that. So
target_colsnames the task ID columns whosecombinations enumerate those targets, and the allocation unit is derived as
model_idplus every remaining task ID column. One problem is solved percombination.
That mirrors how a hubverse
compound_taskid_setworks: the columns you namevary within a group, the rest are held constant. With
target_cols = "location"on thehubExamplesforecast data you get 24allocation problems (3 models x 2 reference dates x 4 horizons), each pooling 2
locations.
This is the part most worth review attention — it is an interpretation of the
hubverse data model, not a mechanical port.
Functions
as_alloscore_df()— exported, because it is where a mis-specified join ortarget set becomes visible. Returns one row per allocation unit with a
forecastslist column holding each target's quantiles, itsdistfromq-builtcdf and quantile function, and the observed outcome.
allocate_model_out()— stops after the allocation, for when the question iswhat a forecast implies you should do rather than how good it was.
alloscore_model_out()— allocates and scores, withsummarize/byin theshape
hubEvals::score_model_out()uses.Validation
Follows
hubEvals:get_task_id_cols()andvalidate_model_oracle_out()mirror its helpers, and a single supported
output_typeis enforced up front.Only
quantileis accepted;sample,cdf,pmf,meanandmedianerrorwith an explanation rather than silently doing something wrong.
Rejected inputs covered by tests: several output types at once, a non-quantile
output type, non-numeric quantile levels,
oracle_outputwith unexpectedcolumns or missing coverage, duplicated forecasts, and
target_colsnamingsomething that is not a task ID.
Tests
418 total, up from 306, in two commits.
The interface adds a test worth looking at,
"alloscore_model_out matches the core API on a single unit": it routes one
allocation problem through both the hubverse entry point and the core
alloscore()and asserts the scores agree, which is what makes this a refactorof the interface rather than a second implementation.
The equivalence fixtures then prove both agree with the original package.
The original has no test suite, so there is no pre-existing ground truth to
port:
data-raw/make_legacy_reference.Rruns the original code once over afixed set of fixtures and freezes the results under
tests/testthat/testdata/, andtest-legacy-equivalence.Rchecks this packagereproduces them. The original therefore stays out of this package's
dependencies while the numerics stay pinned, so any future change in behaviour
has to be deliberate.
legacy_hub_quantile_scores.csvis the one that matters here: hubverse exampledata pushed through the original's hand-rolled pipeline, which
alloscore_model_out()must reproduce. It ships alongside the function ratherthan in a later PR, so the interface does not land unproven.
legacy_reference_meta.jsonrecords the original's git SHA (0477ae9) and theR,
distfromqandhubExamplesversions used. Tolerances are1e-6onallocations and
1e-8on scores. A mutation check confirms the suite is notvacuous: widening
eps_Kfrom 1% to 5% fails these tests and the ZXH tests,and restoring it passes all 54.
Verified locally
R CMD check: Status: OK. 418 tests, 0 failures, 0 warnings, 0 skips.lintr: 0 lints.air format . --check: clean.