Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,7 @@ The catalog-side orphan-recovery subcommands (`dedup-deletions`, `find-orphans`,

`drop-partitions` and `list-droppable-partitions` implement server-side partition retention (megaduck `events_nrt` 14-day campaign and the daily cron). They forge an engine whole-file-drop commit directly in the catalog (snapshot row with `next_file_id = baseline + 1` — the F1 stats-ObjectCache bust, per-row re-verified `end_snapshot` UPDATE, RELATIVE stats decrement from `UPDATE ... RETURNING` sums, `deleted_from_table:<id>` conflict token) because the engine's DELETE cannot complete against the nrt commit rate (insert-vs-delete OCC conflict is a hard abort). These two ops (and `drop-orphan-inline-tables`, below) are the ONLY ones that bypass the duckdb-attach `postgres_execute` path — they use a direct psycopg (libpq) connection (READ COMMITTED, real RETURNING, no wedged-duckdb teardown core dump), so `main()` skips `connect()` for them (`_DIRECT_PG_COMMANDS`). Dry-run is the default; `--execute` writes. Safety posture: leaf = one fully-resolved partition tuple = one short transaction = one snapshot; per-leaf advisory-lock re-acquisition (`pg_advisory_unlock_all` + `pg_try_advisory_lock`, refcount-safe); fresh viaduck cursor guard per leaf read from the durable `viaduck.viaduck_state` store (fail-closed on missing/empty, run-abort on consumer lag past cutoff+48h, `--no-viaduck-guard` prints what it overrides); structural floor at selection (cutoff+48h); partition-spec pinning by live `partition_id` with fail-loud aborts on old-spec/unageable/rot files (remedy: `repair-partition-values`). The stats assertion refuses to take `record_count` to 0, so a table's FINAL leaf is undroppable by design (the engine has a distinct empty-table path we deliberately don't emulate). Full design and hazard list: `DROP_PARTITIONS_PLAN.md` and `partition-drop-findings.md`.

`drop-orphan-inline-tables` drops the Postgres tables behind DuckLake data inlining (`public.ducklake_inlined_data_<table_id>_<schema_version>`, registered in `ducklake_inlined_data_tables`) whose DuckLake table no retained snapshot can reach, and deletes their registry rows. DuckLake's GC (`DropEmptySupersededInlinedTables`) only handles superseded-and-empty tables, so DROP+CREATE churn can leak one Postgres table per dropped DuckLake table; every leaked table adds rows to `pg_class` / `pg_attribute` / `pg_type`, and DuckDB's DuckLake ATTACH scans those system catalogs in full — at ~230k leaked tables the ATTACH itself OOMs. So this op is in `_DIRECT_PG_COMMANDS` and must never need the ATTACH. The orphan predicate (`_inline_orphan_predicate`) is the range-overlap predicate of the `ducklake_unreachable_inline_tables` metric; keep them in lockstep (an integration test runs the metric's SQL against the same catalog and asserts equal counts). Per batch (`--batch-size`, default 500, absolute ceiling 2000, but the EFFECTIVE size is derived from the server at run time: a quarter of `max_locks_per_transaction * (max_connections + max_prepared_transactions)` — the lock table's documented FLOOR, not its total — divided by the 5 shared lock-table entries one DROP of an inlined table holds until commit, measured off `pg_locks` on PG 16/17/18 as heap + TOAST heap + TOAST index + 2 `pg_type` object locks from `AcquireDeletionLock`. A stock/CNPG-default server (64 × 100) caps at 320, below the default; RDS's default `max_connections` (~1,800 on db.r7g.large, 5,000 on 2xlarge) puts the raw cap in the thousands, so on the production targets the 2000 ceiling governs and the derivation changes nothing. SQLSTATE `53200` (generic `out_of_memory`, not lock-table-specific) is retryable and halves the size for the rest of the run, re-growing after 10 clean batches; the 10-row floor is reachable only across batches because `_DROP_INLINE_MAX_ATTEMPTS` = 5 caps one batch's chain at 500/250/125/62/31 and then exits FATALLY on purpose — a lock table still full after that is a catalog problem a cron must not paper over) ONE transaction selects the next keyset page of orphans (`FOR UPDATE OF idt SKIP LOCKED`), then per present table takes `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE`, checks emptiness with `SELECT 1 ... LIMIT 1` (never `reltuples`, which is a stale estimate), skips and WARNs on non-empty tables and on names that do not fullmatch `ducklake_inlined_data_<n>_<n>`, `DROP TABLE IF EXISTS`, and deletes the matching registry rows. The predicate is re-evaluated per batch. Execute mode takes the maintenance advisory lock once (`_lock_refresh`). It never runs `VACUUM FULL` (ACCESS EXCLUSIVE on system catalogs); it logs a hint instead. `--run-budget-s` bounds the wall clock, because the run holds the maintenance advisory lock throughout and `expire` / `cleanup` / `compaction` queue behind it — the tenant maintenance CronJob runs ~6 recipes under ONE `activeDeadlineSeconds` (default 18000, `$cActiveDeadline` in charts `crossplane-config/templates/composition.yaml`), so the chain-safe `drop-orphan-inline-tables-default` wrapper passes 1800s. It is enforced between batches AND inside one (a deadline passed into `_drop_inline_batch`: `statement_timeout` is per statement, a batch issues ~1 + 3×batch + 1 of them, and PG 16 has no `transaction_timeout`; on the deadline the batch takes no new row, finishes the one in flight, commits its registry deletes and returns `truncated`), and the retry backoff checks it instead of sleeping budget away. A budgeted stop logs `budget_exhausted=true` and exits 0, and converges on the next tick because keyset paging restarts from the predicate, not from a saved cursor — but only for DROPPABLE rows: rows skipped as non-empty or as unexpected names are never deleted, so every run re-selects them from the head of the registry and spends batch capacity on them (the done line logs `skipped_share`). Gauges `maintenance_inline_orphans_{dropped_total,skipped_nonempty_total,remaining,budget_exhausted}` are unlabeled and therefore registered only for an execute run of this command, so other crons' pushes cannot zero them. Tests: `tests/integration/test_drop_orphan_inline_tables.py` runs against a throwaway real Postgres (`MILLPOND_TEST_PG_DSN`, else local `initdb`/`pg_ctl`, else docker).
`drop-orphan-inline-tables` drops the Postgres tables behind DuckLake data inlining (`public.ducklake_inlined_data_<table_id>_<schema_version>`, registered in `ducklake_inlined_data_tables`) whose DuckLake table no retained snapshot can reach, and deletes their registry rows. DuckLake's GC (`DropEmptySupersededInlinedTables`) only handles superseded-and-empty tables, so DROP+CREATE churn can leak one Postgres table per dropped DuckLake table; every leaked table adds rows to `pg_class` / `pg_attribute` / `pg_type`, and DuckDB's DuckLake ATTACH scans those system catalogs in full — at ~230k leaked tables the ATTACH itself OOMs. So this op is in `_DIRECT_PG_COMMANDS` and must never need the ATTACH. The orphan predicate (`_inline_orphan_predicate`) is the range-overlap predicate of the `ducklake_unreachable_inline_tables` metric; keep them in lockstep (an integration test runs the metric's SQL against the same catalog and asserts equal counts). Per batch (`--batch-size`, default 500, absolute ceiling 2000, but the EFFECTIVE size is derived from the server at run time: a quarter of `max_locks_per_transaction * (max_connections + max_prepared_transactions)` — the lock table's documented FLOOR, not its total — divided by the 5 shared lock-table entries one DROP of an inlined table holds until commit, measured off `pg_locks` on PG 16/17/18 as heap + TOAST heap + TOAST index + 2 `pg_type` object locks from `AcquireDeletionLock`. A stock/CNPG-default server (64 × 100) caps at 320, below the default; RDS's default `max_connections` (~1,800 on db.r7g.large, 5,000 on 2xlarge) puts the raw cap in the thousands, so on the production targets the 2000 ceiling governs and the derivation changes nothing. SQLSTATE `53200` (generic `out_of_memory`, not lock-table-specific) is retryable and halves the size for the rest of the run, re-growing after 10 clean batches; the 10-row floor is reachable only across batches because `_DROP_INLINE_MAX_ATTEMPTS` = 5 caps one batch's chain at 500/250/125/62/31 and then exits FATALLY on purpose — a lock table still full after that is a catalog problem a cron must not paper over) ONE transaction selects the next keyset page of orphans (`FOR UPDATE OF idt SKIP LOCKED`), then per present table takes `LOCK TABLE ... IN ACCESS EXCLUSIVE MODE`, checks emptiness with `SELECT 1 ... LIMIT 1` (never `reltuples`, which is a stale estimate), skips and WARNs on non-empty tables and on names that do not fullmatch `ducklake_inlined_data_<n>_<n>`, `DROP TABLE IF EXISTS`, and deletes the matching registry rows. The predicate is re-evaluated per batch. Execute mode takes the maintenance advisory lock once (`_lock_refresh`). It never runs `VACUUM FULL` (ACCESS EXCLUSIVE on system catalogs); it logs a hint instead. `--run-budget-s` bounds the wall clock, because the run holds the maintenance advisory lock throughout and `expire` / `cleanup` / `compaction` queue behind it — the bound that matters is the cron cadence, not the chain's `activeDeadlineSeconds` (18000 by default): three prod-us tenants run the whole chain every 15 minutes under `concurrencyPolicy: Forbid`, so a step that outlives a tick skips the next tick for every other step. The chain-safe `drop-orphan-inline-tables-default` wrapper reads the budget from `MILLPOND_INLINE_RUN_BUDGET_S` (default 300s) so a tenant can be raised from its CronJob spec without touching the chain args (a parameterized recipe in a `just` chain eats the next word). It is enforced between batches AND inside one (a deadline passed into `_drop_inline_batch`: `statement_timeout` is per statement, a batch issues ~1 + 3×batch + 1 of them, and PG 16 has no `transaction_timeout`; on the deadline the batch takes no new row, finishes the one in flight, commits its registry deletes and returns `truncated`), and the retry backoff checks it instead of sleeping budget away. A budgeted stop logs `budget_exhausted=true` and exits 0, and converges on the next tick because keyset paging restarts from the predicate, not from a saved cursor — but only for DROPPABLE rows: rows skipped as non-empty or as unexpected names are never deleted, so every run re-selects them from the head of the registry and spends batch capacity on them (the done line logs `skipped_share`). Gauges `maintenance_inline_orphans_{dropped_total,skipped_nonempty_total,remaining,budget_exhausted}` are unlabeled and therefore registered only for an execute run of this command, so other crons' pushes cannot zero them. Tests: `tests/integration/test_drop_orphan_inline_tables.py` runs against a throwaway real Postgres (`MILLPOND_TEST_PG_DSN`, else local `initdb`/`pg_ctl`, else docker).

`cleanup` and `cleanup-all` log a single structured throughput line on completion: `cleanup throughput: files_processed=N elapsed_s=T rate_obj_s=R queue_depth_after=A`. `files_processed` comes directly from `len(result)` — the count of rows `ducklake_cleanup_old_files` returned — rather than a queue-depth delta. A delta would be misleading whenever any other writer enqueues deletions during the call (the maintenance advisory lock by design only mutexes maintenance invocations, not arbitrary writers); `len(result)` is accurate regardless. `queue_depth_after` is queried with a single post-call snapshot and shows remaining work but is not used in the rate calculation. The line is intentionally suppressed on `cleanup --dry-run` because dry-run returns preview rows and a rate computed from those would falsely claim work was done. `cleanup-all` has no dry-run form at all — the CLI rejects `--dry-run` rather than silently no-opping (preview via `cleanup-dry-run` / `fsck-dry-run`).

Expand Down
12 changes: 11 additions & 1 deletion tests/integration/test_drop_orphan_inline_tables.py
Original file line number Diff line number Diff line change
Expand Up @@ -589,7 +589,17 @@ def test_default_wrapper_renders_the_run_budget(self):
rendered = out.stdout + out.stderr
assert out.returncode == 0, rendered
assert "drop-orphan-inline-tables --batch-size '500'" in rendered
assert "--run-budget-s '1800'" in rendered
assert "--run-budget-s '300'" in rendered
# The budget is an environment knob, so a tenant can be raised from
# the CronJob spec without touching the chain args.
raised = subprocess.run(
[just, "--justfile", justfile, "--dry-run", "drop-orphan-inline-tables-default"],
capture_output=True,
text=True,
env={**os.environ, "MILLPOND_INLINE_RUN_BUDGET_S": "900"},
)
assert raised.returncode == 0, raised.stdout + raised.stderr
assert "--run-budget-s '900'" in raised.stdout + raised.stderr
assert "--max-batches" not in rendered

def test_empty_run_budget_renders_an_unbudgeted_command(self):
Expand Down
42 changes: 24 additions & 18 deletions tools/justfile
Original file line number Diff line number Diff line change
Expand Up @@ -204,19 +204,21 @@ drop-orphan-inline-tables-dry-run batch_size="500":
# purpose.
#
# run_budget_s bounds the wall clock because the whole run holds the
# maintenance advisory lock, which expire / cleanup / compaction wait on: the
# tenant maintenance CronJob runs ~6 recipes under ONE activeDeadlineSeconds
# (default 18000, `$cActiveDeadline ... default 18000` in the charts
# crossplane-config/templates/composition.yaml), so 18000 / 6 = 3000s per
# recipe and 1800s leaves headroom for the others. Stopping on the budget is
# a clean exit 0, not a failure: keyset paging restarts from the orphan
# predicate rather than a saved cursor, so the next cron tick resumes the
# backlog of droppable rows. Pass an empty run_budget_s to drain a backlog
# unbudgeted by hand.
# maintenance advisory lock, which expire / cleanup / compaction wait on. The
# bound that matters is the cron CADENCE, not the activeDeadlineSeconds
# (18000 by default): three prod-us tenants run the whole chain every 15
# minutes under concurrencyPolicy Forbid, so a step that holds the chain for
# longer than a tick skips the next tick for every other step. 300s keeps
# the chain inside one tick; the chain wrapper below reads it from
# MILLPOND_INLINE_RUN_BUDGET_S so a tenant can be raised without touching
# the chain args. Stopping on the budget is a clean exit 0, not a failure:
# keyset paging restarts from the orphan predicate rather than a saved
# cursor, so the next cron tick resumes the backlog of droppable rows. Pass
# an empty run_budget_s to drain a backlog unbudgeted by hand.

# Drop inlined-data tables of dropped DuckLake tables (direct Postgres; works when ATTACH cannot)
[group('lifecycle')]
drop-orphan-inline-tables batch_size="500" max_batches="" run_budget_s="1800": _confirm-target
drop-orphan-inline-tables batch_size="500" max_batches="" run_budget_s="300": _confirm-target
env {{ _target_env }} python {{ ducklake_maintenance }} drop-orphan-inline-tables --batch-size {{ quote(batch_size) }} {{ if run_budget_s != "" { "--run-budget-s " + quote(run_budget_s) } else { "" } }} {{ if max_batches != "" { "--max-batches " + quote(max_batches) } else { "" } }}

# REMOVE-WHEN: upstream-ducklake-fix-deployed
Expand Down Expand Up @@ -353,14 +355,18 @@ compact-all-tiers-default: (compact-all-tiers)
# snapshot expiry at the 7-day fleet default, chain-safe
expire-7d: (expire "7")

# The 1800s run budget is passed explicitly (not left to the recipe default)
# because this is the form the maintenance cron runs: one of ~6 recipes under a
# single 18000s activeDeadlineSeconds, so 18000 / 6 = 3000s each and 1800s
# leaves headroom. A budgeted stop is not a failure - the next tick resumes
# from the orphan predicate (keyset paging, no saved cursor).

# drop-orphan-inline-tables at the default batch size, no batch limit, 30m budget, chain-safe
drop-orphan-inline-tables-default: (drop-orphan-inline-tables "500" "" "1800")
# The run budget comes from MILLPOND_INLINE_RUN_BUDGET_S (default 300s), not
# from the recipe default, because this is the form the maintenance cron runs
# and the cron's cadence is the real bound: three prod-us tenants run the whole
# chain every 15 minutes under concurrencyPolicy Forbid, so a step that holds
# the chain for 30 minutes skips their ticks (compaction included). 300s keeps
# the chain inside one tick with the other steps. A budgeted stop is not a
# failure - the next tick resumes from the orphan predicate (keyset paging, no
# saved cursor). Raise it per tenant with the env var, never by editing the
# chain args: a parameterized recipe in a just chain eats the next word.

# drop-orphan-inline-tables at the default batch size, no batch limit, env budget (default 300s), chain-safe
drop-orphan-inline-tables-default: (drop-orphan-inline-tables "500" "" env_var_or_default("MILLPOND_INLINE_RUN_BUDGET_S", "300"))

# Run all tiered compactions sequentially (tier 1 -> 2 -> 3) then delete scheduled S3 files
[group('compaction')]
Expand Down
Loading