| title | Quickstart |
|---|---|
| description | Run a complete robotics data pipeline with Refiner |
Refiner provides open-source data-processing utilities and reference implementations from our published research. Contact us directly for access to our proprietary models and pipelines.
Refiner is an open-source library for building robotics data pipelines. A pipeline describes how to read data, transform rows and episodes, and write the result. You can process a wide range of formats out of the box, use models for labeling and scoring, and inspect pipelines locally while you develop. When you want parallel execution on your machine, run the same code with local workers.
Readers and writers are sharded, so most pipelines do not need to download or materialize the entire dataset before doing useful work. Refiner streams shards through the pipeline and writes outputs as they are produced.
Install the Refiner package in the Python environment where you want to build or run the pipeline:
pip install macrodata-refiner[hf,video]The hf and video extras are optional, but the example below uses them for
Hugging Face paths and
video data. To install every optional
dependency, use pip install macrodata-refiner[all].
Most Refiner pipelines follow the same shape:
read data → transform → write result
For example:
import refiner as mdr
pipeline = (
mdr.read_lerobot("hf://datasets/macrodata/aloha_static_battery_ep005_009")
.map(lambda row: row.update(task="battery insertion"))
.write_lerobot("hf://buckets/macrodata/test_bucket/aloha_static_with_task")
)This example reads a small Hugging Face dataset, adds a task label to each episode, and writes the result back to a Hugging Face bucket.
Each step returns a new pipeline value:
| Step | What it does |
|---|---|
read_lerobot(...) |
Reads robotics episodes with the LeRobot reader. |
.map(...) |
Updates each row or episode with a task label. |
.write_lerobot(...) |
Writes the transformed dataset with the LeRobot writer. |
Nothing runs when you create the pipeline. Refiner executes it only when you
inspect rows with methods like take() or launch
local workers.
Start by inspecting a small amount of data in process. This block is self-contained and inspects the pipeline before the writer:
import refiner as mdr
pipeline = (
mdr.read_lerobot("hf://datasets/macrodata/aloha_static_battery_ep005_009")
.map(lambda row: row.update(task="battery insertion"))
)
row = pipeline.take(1)
print(row[0])Output:
LeRobotRow
episode_id: '0'
num_frames: 600
task: 'battery insertion'
fps: 50
robot_type: 'unknown'
frame_data (row.to_frame_table()):
actions (row.actions): float[600, 14]
states (row.states): float[600, 14]
timestamps (row.timestamps): float[600]
videos (row.videos):
observation.images.cam_high: video[480, 640, 3]@50fps
observation.images.cam_left_wrist: video[480, 640, 3]@50fps
observation.images.cam_low: video[480, 640, 3]@50fps
observation.images.cam_right_wrist: video[480, 640, 3]@50fps
stats: ['action', 'episode_index', 'frame_index', 'index', 'next.done', 'observation.images.cam_high', 'observation.images.cam_left_wrist', 'observation.images.cam_low', '... +4 more']
take() executes lazily and stops after the requested number of rows, so it is
the fastest way to check schemas,
media references, and transform outputs
before launching a full job.
Consider this pipeline:
import refiner as mdr
def log_stats(row):
row.log_histogram("frames", row.num_frames, unit="frames", per="episode")
mdr.logger.info(
"episode={} task={!r} frames={} cameras={}",
row.episode_id,
row.task,
row.num_frames,
sorted(row.videos),
)
return row
pipeline = (
mdr.read_lerobot("hf://datasets/macrodata/aloha_static_battery_ep005_009")
.map(log_stats)
)
pipeline.launch_local(name="quickstart-aloha-summary")Local launch is the standard way to run a complete Refiner pipeline. It distributes shards across worker processes on your machine.
Here is a more elaborate example, reading from and writing to Hugging Face buckets.
import refiner as mdr
dataset = "hf://datasets/yifengzhu-hf/LIBERO-datasets/libero_spatial"
# Replace this with your output bucket (HF, S3, GCP, etc).
output = "hf://buckets/macrodata/test_bucket/libero-spatial"
(
mdr.read_hdf5(
dataset,
groups="/data/demo_*",
datasets={
"raw_action": "actions",
"observation.images.image": "obs/agentview_rgb",
"observation.images.wrist_image": "obs/eye_in_hand_rgb",
"ee_state": "obs/ee_states",
"gripper_state": "obs/gripper_states",
},
file_path_column="file_path",
cache_remote_files=True,
)
.map(
lambda row: row.update(
task=str(row["file_path"])
.rsplit("/", 1)[-1]
.removesuffix(".hdf5")
.removesuffix("_demo")
.replace("_", " ")
)
)
.to_robot_rows(
task_key="task",
fps=10.0,
robot_type="libero",
action_key="raw_action",
state_key=("ee_state", "gripper_state"),
video_keys={
"observation.images.image": "observation.images.image",
"observation.images.wrist_image": "observation.images.wrist_image",
},
)
.write_lerobot(output, max_video_prepare_in_flight=2)
.launch_local(
name="libero-spatial-subset",
num_workers=2,
)
)This example converts the public LIBERO spatial HDF5 subset to LeRobot using local workers. It reads one demo group per row, derives the task label from the filename, turns action/state/image arrays into robotics episodes, encodes the two camera streams as videos, and writes a LeRobot dataset to your output bucket.
The input dataset is public, but writing to your Hugging Face bucket requires an
HF_TOKEN in the environment inherited by the local workers:
export HF_TOKEN="your-write-token"For the full four-suite LIBERO conversion, see Libero HDF5.
- Learn the execution options in Running Pipelines.
- Learn readers in Reading Data.
- Learn the episode model in Episode Data.
- Learn common row operations in Transforms.
- Learn LeRobot output details in Writing LeRobot.