Version or commit
94a78fde8c5e49bbbdb4069f823e845ca3ee0d42 (current main)
Environment
macOS / Linux, Python 3.11+, DuckDB 1.4+
Minimal reproduction
import tempfile
from pathlib import Path
import pyarrow as pa
import pyarrow.parquet as pq
from hflow import ManifestSplitSettings, split_manifest
with tempfile.TemporaryDirectory() as td:
td = Path(td)
source = td / "ambiguous.parquet"
table = pa.table({
"sample_id": ["s1", "s2", "s3"],
"group": ["g1", "g2", "g3"],
"metadata": ["m1", "m2", "m3"],
"METADATA": ["M1", "M2", "M3"],
})
pq.write_table(table, source)
out = td / "split"
split_manifest(source, out, settings=ManifestSplitSettings("sample_id", ("group",)))
written_table = pq.read_table(out / "train.parquet")
print("Original columns:", table.column_names)
print("Written columns: ", written_table.column_names)
assert written_table.column_names == table.column_names
Expected behavior
Either:
split_manifest partitions the file preserving original column names without modification, OR
- If ambiguous case-colliding column names are present,
split_manifest refuses to split and raises ValueError("manifest column names must be distinct ignoring case") (matching the pattern in manifest_deduplication.py:143 to avoid silent column renaming by DuckDB).
Additionally, ManifestSplitSettings.__post_init__ should reject case-colliding group_columns (e.g. ("group", "GROUP")), matching ManifestDeduplicationSettings.
Actual behavior
DuckDB silently renames METADATA to METADATA_1 when creating the source_manifest view:
Original columns: ['sample_id', 'group', 'metadata', 'METADATA']
Written columns: ['sample_id', 'group', 'metadata', 'METADATA_1']
AssertionError
The output Parquet files (train.parquet, development.parquet, test.parquet) have their schemas silently corrupted, violating split_manifest's contract ("Partition a local Parquet file, preserving every row and column.").
Furthermore, if a case-colliding column is specified in group_columns, _read_sample_identities queries DESCRIBE source_manifest and raises ValueError: manifest is missing columns: ['METADATA'], even though METADATA is present in the source manifest.
Additional context
In commit 0c4583b (#658), manifest_deduplication.py addressed this exact DuckDB renaming behavior by inspecting parquet_schema(?) before creating the view:
# Check original top-level names before DuckDB can rename ambiguous fields.
source_columns: list[str] = []
nested_fields_remaining = 0
schema_fields = connection.execute(
"SELECT name, num_children FROM parquet_schema(?)", [str(input_snapshot)]
).fetchall()
for column_name, child_count in schema_fields[1:]:
if nested_fields_remaining:
nested_fields_remaining += (child_count or 0) - 1
else:
source_columns.append(column_name)
nested_fields_remaining = child_count or 0
if len({column.casefold() for column in source_columns}) != len(source_columns):
raise ValueError("manifest column names must be distinct ignoring case")
and added test_ambiguous_source_names_cannot_be_silently_renamed in tests/test_manifest_deduplication.py.
manifest_splits.py was created earlier and did not include this check.
Version or commit
94a78fde8c5e49bbbdb4069f823e845ca3ee0d42(currentmain)Environment
macOS / Linux, Python 3.11+, DuckDB 1.4+
Minimal reproduction
Expected behavior
Either:
split_manifestpartitions the file preserving original column names without modification, ORsplit_manifestrefuses to split and raisesValueError("manifest column names must be distinct ignoring case")(matching the pattern inmanifest_deduplication.py:143to avoid silent column renaming by DuckDB).Additionally,
ManifestSplitSettings.__post_init__should reject case-collidinggroup_columns(e.g.("group", "GROUP")), matchingManifestDeduplicationSettings.Actual behavior
DuckDB silently renames
METADATAtoMETADATA_1when creating thesource_manifestview:The output Parquet files (
train.parquet,development.parquet,test.parquet) have their schemas silently corrupted, violatingsplit_manifest's contract ("Partition a local Parquet file, preserving every row and column.").Furthermore, if a case-colliding column is specified in
group_columns,_read_sample_identitiesqueriesDESCRIBE source_manifestand raisesValueError: manifest is missing columns: ['METADATA'], even thoughMETADATAis present in the source manifest.Additional context
In commit
0c4583b(#658),manifest_deduplication.pyaddressed this exact DuckDB renaming behavior by inspectingparquet_schema(?)before creating the view:and added
test_ambiguous_source_names_cannot_be_silently_renamedintests/test_manifest_deduplication.py.manifest_splits.pywas created earlier and did not include this check.