Skip to content

Make the PostgreSQL batch benchmark runner turnkey - #56

Closed
estebanzimanyi wants to merge 2 commits into
MobilityDB:masterfrom
estebanzimanyi:feat/batch-mbdb-turnkey
Closed

estebanzimanyi wants to merge 2 commits into
MobilityDB:masterfrom
estebanzimanyi:feat/batch-mbdb-turnkey

Conversation

@estebanzimanyi

Copy link
Copy Markdown
Member

The batch runner bench_mbdb.sh reads a load_mbdb.sql loader and per-query qNN.sql files that are absent from the repository, so the MobilityDB side of the cross-platform batch benchmark cannot run.

This makes the PostgreSQL side turnkey:

  • bench/load_mbdb.sql loads the cross-platform portability export (vehicles.csv / trips.csv / query_*.csv from berlinmod_portability_export) into the schema queries.sql expects — Trips and Vehicles plus the Query* parameter selectors, the trip_h3 prefilter column, and the GiST / SP-GiST / th3 indexes the tier logic toggles. The trip trajectory is parsed from its hex-EWKB; geomWKT is the SRID-tagged EWKT text.
  • bench_mbdb.sh reads each query from the shared queries.sql, split on its -- BerlinMOD Qn: markers, and surfaces query failures instead of recording phantom sub-second timings.
  • setup/generate_data.sh orchestrates berlinmod_generate and berlinmod_portability_export to produce the CSV corpus (default scale factor 0.005), closing the reference in bench.sh.

At scale factor 0.005 (1620 trips) all eighteen queries (q01–q17 + qrt) load and run under tier 3.

The batch runner bench_mbdb.sh reads a load_mbdb.sql loader and per-query
qNN.sql files that are absent from the repository, so the MobilityDB side of
the cross-platform benchmark cannot run.

Add bench/load_mbdb.sql, which loads the cross-platform portability export
(vehicles/trips/query_*.csv produced by berlinmod_portability_export) into the
schema the shared queries.sql expects: Trips and Vehicles plus the Query*
parameter selectors, the trip_h3 prefilter column, and the GiST/SP-GiST/th3
indexes the tier logic toggles. The trip trajectory is parsed from its
hex-EWKB, and geomWKT is carried as the SRID-tagged EWKT text.

Rewire bench_mbdb.sh to read each query from the shared queries.sql, split on
its "-- BerlinMOD Qn:" markers, and surface query failures instead of
recording phantom sub-second timings.

Add setup/generate_data.sh, orchestrating berlinmod_generate and
berlinmod_portability_export to produce the CSV corpus (default scale
factor 0.005), closing the reference in bench.sh.
The Deploy documentation job is gated to master, so on every pull request it
reports as a skipped check. Fold its gh-pages publish into the Generate
documentation job as conditional steps guarded on a master push: a skipped
step emits no check, so pull requests show a single always-green
documentation check and no skipped deploy job. Deploy behaviour is unchanged.
@estebanzimanyi

Copy link
Copy Markdown
Member Author

Closing per maintainer direction: this duplicates the canonical loading path. The repository's own berlinmod_load.sql is the canonical PostgreSQL BerlinMOD loader; the batch runner should build on that rather than a separate hand-authored one over the portability-export CSVs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant