Skip to content

[Bug]: split_manifest silently corrupts output schema when source parquet contains case-colliding column names #683

Description

@shobhitagnihotri69

Version or commit

94a78fde8c5e49bbbdb4069f823e845ca3ee0d42 (current main)

Environment

macOS / Linux, Python 3.11+, DuckDB 1.4+

Minimal reproduction

import tempfile
from pathlib import Path
import pyarrow as pa
import pyarrow.parquet as pq
from hflow import ManifestSplitSettings, split_manifest

with tempfile.TemporaryDirectory() as td:
    td = Path(td)
    source = td / "ambiguous.parquet"
    table = pa.table({
        "sample_id": ["s1", "s2", "s3"],
        "group": ["g1", "g2", "g3"],
        "metadata": ["m1", "m2", "m3"],
        "METADATA": ["M1", "M2", "M3"],
    })
    pq.write_table(table, source)
    out = td / "split"
    split_manifest(source, out, settings=ManifestSplitSettings("sample_id", ("group",)))
    written_table = pq.read_table(out / "train.parquet")
    print("Original columns:", table.column_names)
    print("Written columns: ", written_table.column_names)
    assert written_table.column_names == table.column_names

Expected behavior

Either:

  1. split_manifest partitions the file preserving original column names without modification, OR
  2. If ambiguous case-colliding column names are present, split_manifest refuses to split and raises ValueError("manifest column names must be distinct ignoring case") (matching the pattern in manifest_deduplication.py:143 to avoid silent column renaming by DuckDB).

Additionally, ManifestSplitSettings.__post_init__ should reject case-colliding group_columns (e.g. ("group", "GROUP")), matching ManifestDeduplicationSettings.

Actual behavior

DuckDB silently renames METADATA to METADATA_1 when creating the source_manifest view:

Original columns: ['sample_id', 'group', 'metadata', 'METADATA']
Written columns:  ['sample_id', 'group', 'metadata', 'METADATA_1']
AssertionError

The output Parquet files (train.parquet, development.parquet, test.parquet) have their schemas silently corrupted, violating split_manifest's contract ("Partition a local Parquet file, preserving every row and column.").

Furthermore, if a case-colliding column is specified in group_columns, _read_sample_identities queries DESCRIBE source_manifest and raises ValueError: manifest is missing columns: ['METADATA'], even though METADATA is present in the source manifest.

Additional context

In commit 0c4583b (#658), manifest_deduplication.py addressed this exact DuckDB renaming behavior by inspecting parquet_schema(?) before creating the view:

# Check original top-level names before DuckDB can rename ambiguous fields.
source_columns: list[str] = []
nested_fields_remaining = 0
schema_fields = connection.execute(
    "SELECT name, num_children FROM parquet_schema(?)", [str(input_snapshot)]
).fetchall()
for column_name, child_count in schema_fields[1:]:
    if nested_fields_remaining:
        nested_fields_remaining += (child_count or 0) - 1
    else:
        source_columns.append(column_name)
        nested_fields_remaining = child_count or 0
if len({column.casefold() for column in source_columns}) != len(source_columns):
    raise ValueError("manifest column names must be distinct ignoring case")

and added test_ambiguous_source_names_cannot_be_silently_renamed in tests/test_manifest_deduplication.py.

manifest_splits.py was created earlier and did not include this check.

Activity

  1. added a commit that references this issue on Oct 5, 2026
    04fa15c
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions