Skip to content

fix(episode): allow null elements in numeric list columns in to_arrow (#662) - #664

Merged
kstonekuan merged 1 commit into
Hebbian-Robotics:mainfrom
shobhitagnihotri69:fix/episode-to-arrow-nullable-numeric-sequence
Oct 3, 2026
Merged

kstonekuan merged 1 commit into
Hebbian-Robotics:mainfrom
shobhitagnihotri69:fix/episode-to-arrow-nullable-numeric-sequence

Conversation

@shobhitagnihotri69

Copy link
Copy Markdown
Contributor

Summary

Fixes an issue where ChannelData.to_arrow() crashed with ValueError: field mixes nested values with primitive values or silently omitted the column when a numeric list began with a None/null element (e.g., [null, 2.0]).

Root Cause

_is_numeric_sequence previously inspected only value[0]. If a list or tuple had a leading None, _is_numeric_scalar(value[0]) evaluated to False. Because it was neither an Arrow scalar nor an empty sequence, _arrow_column_values treated it as a nested structure (saw_nested = True), raising a field conflict when other messages contained numbers or omitting the field if all messages began with null.

Changes

Per maintainer guidance in #662:

  • _is_numeric_sequence: Check the first non-null element instead of strictly value[0].
  • _is_empty_sequence: Treat all-null sequences (e.g. [None, None]) like empty sequences ([]), avoiding false nested flags and keeping the slot untyped unless a typed sample arrives.
  • tests/test_episode.py: Added regression test test_channel_to_arrow_numeric_list_with_null_elements verifying both orders ([1.0, 2.0, None] and [None, 2.0, 3.0]), all-null elements alongside numeric samples, and ensuring all-null list fields are omitted cleanly like empty lists.

Validation

  • uv run pytest tests/test_episode.py passes (10/10).
  • Full test suite passes: 2,431 passed, 49 skipped.
  • uv run ruff check & uv run ruff format --check pass with no warnings.

Closes #662

@greptile-apps

greptile-apps Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[Medium risk] Fixes null handling in numeric list column detection.

This PR needs changes before merging because mixed lists can stop Arrow export and nullable lists can break automatic NumPy selection.

Findings

  1. P1 Mixed list stops Arrow export ▶
  2. P1 Automatic NumPy choice fails ▶
  3. P2 Field error loses its name ▶
Summary

This PR fixes ChannelData.to_arrow() so numeric lists with null entries are kept in Arrow columns. Lists made only of nulls stay untyped and are omitted when no typed sample exists.

  • Numeric lists can start with null and still be recognized.
  • The regression test checks nulls in numeric lists and the omission of an all-null field.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Decoded list] --> B{First non-null item}
  B -->|Number| C[Mark numeric]
  C --> D[Arrow column]
  C --> E[NumPy field choice]
  B -->|No item| F[Leave untyped]
Loading

Reviews (1) · Last reviewed commit: "fix(episode): allow null elements in num..."

Comment thread src/hflow/episode.py
Comment on lines +139 to +141
for item in value:
if item is not None:
return _is_numeric_scalar(item)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Mixed list stops Arrow export

A JSON list like [null, 1.0, {"note": "bad"}] is now marked numeric because its first non-null item is a number. If every sample of that field has this shape, to_arrow() used to skip the nested field. Now it sends the mixed list to Arrow, which cannot make a numeric list from it, so the caller loses the whole table. Check the later items before marking the list numeric.

Knowledge Base Used: Episode storage and identity

Comment thread src/hflow/episode.py
Comment on lines +139 to +141
for item in value:
if item is not None:
return _is_numeric_scalar(item)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Automatic NumPy choice fails

_is_numeric_sequence also helps to_numpy() choose a field. When the first message has a nullable position list like [None, 2.0] and a numeric scalar field, the new check makes to_numpy() pick position instead of the scalar. NumPy gives that list an object dtype, which to_numpy() rejects, so a call that could return the scalar now fails. Keep the Arrow change from altering this choice unless NumPy can use the list.

Knowledge Base Used: Episode storage and identity

Comment thread src/hflow/episode.py
Comment on lines +159 to +160
if isinstance(value, (list, tuple)):
return len(value) == 0 or all(item is None for item in value)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Field error loses its name

An all-null list now counts as having no shape during the column check. If another message has a scalar in the same field, to_arrow() skips its error that names the topic and field, then passes both values to Arrow. The caller loses the useful field-specific error, making the bad recording harder to fix. Keep that error for this mixed shape.

Knowledge Base Used: Episode storage and identity

@kstonekuan kstonekuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, merging.

@kstonekuan
kstonekuan merged commit 3b16505 into Hebbian-Robotics:main Oct 3, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: ChannelData.to_arrow() crashes or silently drops column when JSON array contains null elements

2 participants