Skip to content

Make the MobilityDuck batch benchmark runner turnkey - #57

Closed
estebanzimanyi wants to merge 1 commit into
MobilityDB:masterfrom
estebanzimanyi:feat/batch-mduck-turnkey
Closed

estebanzimanyi wants to merge 1 commit into
MobilityDB:masterfrom
estebanzimanyi:feat/batch-mduck-turnkey

Conversation

@estebanzimanyi

Copy link
Copy Markdown
Member

The batch runner bench_mduck.sh reads a load_mduck.sql loader and per-query qNN.sql files that are absent from the repository, so the MobilityDuck side of the cross-platform batch benchmark cannot run.

This makes the MobilityDuck side turnkey, mirroring the MobilityDB runner:

  • bench/load_mduck.sql loads the cross-platform portability export — the same vehicles.csv / trips.csv / query_*.csv the MobilityDB runner consumes — into the schema queries.sql expects: Trips and Vehicles plus the Query* selectors and the trip_h3 prefilter column. SRID is resolved by MEOS parsers, never by hand: the trip via tgeompointFromHexEWKB and the static geometries via ST_GeomFromText (which accepts the SRID-tagged EWKT).
  • bench_mduck.sh reads each query from the shared queries.sql, split on its -- BerlinMOD Qn: markers, and reports any query the loaded extension cannot run instead of recording a phantom sub-second timing.

At scale factor 0.005 (632 vehicles / 1620 trips) the data loads and every query the extension supports runs and times, including the h3 prefilter (geoToH3IndexSet + eEq + eIntersects). A query that needs a function a given extension build does not expose (for example geoToH3Cell or the eDwithinPairs set-set join) is reported and excluded, so the harness is transparent about extension coverage rather than silently green.

The batch runner bench_mduck.sh reads a load_mduck.sql loader and per-query
qNN.sql files that are absent from the repository, so the MobilityDuck side of
the cross-platform benchmark cannot run.

Add bench/load_mduck.sql, which loads the cross-platform portability export
(the same vehicles/trips/query_*.csv the MobilityDB runner consumes) into the
schema queries.sql expects: Trips and Vehicles plus the Query* selectors and
the trip_h3 prefilter. SRID is resolved by MEOS parsers, never by hand: trips
via tgeompointFromHexEWKB, static geometries via ST_GeomFromText (which accepts
the SRID-tagged EWKT). Rewire bench_mduck.sh to read each query from the shared
queries.sql, split on its "-- BerlinMOD Qn:" markers, and report queries the
extension cannot run instead of recording phantom timings.
@estebanzimanyi

Copy link
Copy Markdown
Member Author

Closing per maintainer direction: this duplicates the canonical per-tool loading path. The batch runner should reuse MobilityDuck's own BerlinMOD loader rather than a separate hand-authored one over the portability-export CSVs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant