Repository navigation
fix(episode): allow null elements in numeric list columns in to_arrow (#662) - #664
Conversation
|
| for item in value: | ||
| if item is not None: | ||
| return _is_numeric_scalar(item) |
There was a problem hiding this comment.
A JSON list like [null, 1.0, {"note": "bad"}] is now marked numeric because its first non-null item is a number. If every sample of that field has this shape, to_arrow() used to skip the nested field. Now it sends the mixed list to Arrow, which cannot make a numeric list from it, so the caller loses the whole table. Check the later items before marking the list numeric.
Knowledge Base Used: Episode storage and identity
| for item in value: | ||
| if item is not None: | ||
| return _is_numeric_scalar(item) |
There was a problem hiding this comment.
_is_numeric_sequence also helps to_numpy() choose a field. When the first message has a nullable position list like [None, 2.0] and a numeric scalar field, the new check makes to_numpy() pick position instead of the scalar. NumPy gives that list an object dtype, which to_numpy() rejects, so a call that could return the scalar now fails. Keep the Arrow change from altering this choice unless NumPy can use the list.
Knowledge Base Used: Episode storage and identity
| if isinstance(value, (list, tuple)): | ||
| return len(value) == 0 or all(item is None for item in value) |
There was a problem hiding this comment.
An all-null list now counts as having no shape during the column check. If another message has a scalar in the same field, to_arrow() skips its error that names the topic and field, then passes both values to Arrow. The caller loses the useful field-specific error, making the bad recording harder to fix. Keep that error for this mixed shape.
Knowledge Base Used: Episode storage and identity
Summary
Fixes an issue where
ChannelData.to_arrow()crashed withValueError: field mixes nested values with primitive valuesor silently omitted the column when a numeric list began with aNone/nullelement (e.g.,[null, 2.0]).Root Cause
_is_numeric_sequencepreviously inspected onlyvalue[0]. If a list or tuple had a leadingNone,_is_numeric_scalar(value[0])evaluated toFalse. Because it was neither an Arrow scalar nor an empty sequence,_arrow_column_valuestreated it as a nested structure (saw_nested = True), raising a field conflict when other messages contained numbers or omitting the field if all messages began with null.Changes
Per maintainer guidance in #662:
_is_numeric_sequence: Check the first non-null element instead of strictlyvalue[0]._is_empty_sequence: Treat all-null sequences (e.g.[None, None]) like empty sequences ([]), avoiding false nested flags and keeping the slot untyped unless a typed sample arrives.tests/test_episode.py: Added regression testtest_channel_to_arrow_numeric_list_with_null_elementsverifying both orders ([1.0, 2.0, None]and[None, 2.0, 3.0]), all-null elements alongside numeric samples, and ensuring all-null list fields are omitted cleanly like empty lists.Validation
uv run pytest tests/test_episode.pypasses (10/10).uv run ruff check&uv run ruff format --checkpass with no warnings.Closes #662