Skip to content

Latest commit

 

History

History
235 lines (181 loc) · 7.27 KB

File metadata and controls

235 lines (181 loc) · 7.27 KB
title Quickstart
description Run a complete robotics data pipeline with Refiner

Quickstart

Refiner provides open-source data-processing utilities and reference implementations from our published research. Contact us directly for access to our proprietary models and pipelines.

Refiner is an open-source library for building robotics data pipelines. A pipeline describes how to read data, transform rows and episodes, and write the result. You can process a wide range of formats out of the box, use models for labeling and scoring, and inspect pipelines locally while you develop. When you want parallel execution on your machine, run the same code with local workers.

Readers and writers are sharded, so most pipelines do not need to download or materialize the entire dataset before doing useful work. Refiner streams shards through the pipeline and writes outputs as they are produced.

Install

Install the Refiner package in the Python environment where you want to build or run the pipeline:

pip install macrodata-refiner[hf,video]

The hf and video extras are optional, but the example below uses them for Hugging Face paths and video data. To install every optional dependency, use pip install macrodata-refiner[all].

Example

Most Refiner pipelines follow the same shape:

read data  →  transform  →  write result

For example:

import refiner as mdr

pipeline = (
    mdr.read_lerobot("hf://datasets/macrodata/aloha_static_battery_ep005_009")
    .map(lambda row: row.update(task="battery insertion"))
    .write_lerobot("hf://buckets/macrodata/test_bucket/aloha_static_with_task")
)

This example reads a small Hugging Face dataset, adds a task label to each episode, and writes the result back to a Hugging Face bucket.

Each step returns a new pipeline value:

Step What it does
read_lerobot(...) Reads robotics episodes with the LeRobot reader.
.map(...) Updates each row or episode with a task label.
.write_lerobot(...) Writes the transformed dataset with the LeRobot writer.

Nothing runs when you create the pipeline. Refiner executes it only when you inspect rows with methods like take() or launch local workers.

Inspect a pipeline

Start by inspecting a small amount of data in process. This block is self-contained and inspects the pipeline before the writer:

import refiner as mdr

pipeline = (
    mdr.read_lerobot("hf://datasets/macrodata/aloha_static_battery_ep005_009")
    .map(lambda row: row.update(task="battery insertion"))
)

row = pipeline.take(1)

print(row[0])

Output:

LeRobotRow
  episode_id: '0'
  num_frames: 600
  task: 'battery insertion'
  fps: 50
  robot_type: 'unknown'
  frame_data (row.to_frame_table()):
    actions (row.actions): float[600, 14]
    states (row.states): float[600, 14]
    timestamps (row.timestamps): float[600]
  videos (row.videos):
    observation.images.cam_high: video[480, 640, 3]@50fps
    observation.images.cam_left_wrist: video[480, 640, 3]@50fps
    observation.images.cam_low: video[480, 640, 3]@50fps
    observation.images.cam_right_wrist: video[480, 640, 3]@50fps
  stats: ['action', 'episode_index', 'frame_index', 'index', 'next.done', 'observation.images.cam_high', 'observation.images.cam_left_wrist', 'observation.images.cam_low', '... +4 more']

take() executes lazily and stops after the requested number of rows, so it is the fastest way to check schemas, media references, and transform outputs before launching a full job.

Run locally

Consider this pipeline:

import refiner as mdr

def log_stats(row):
    row.log_histogram("frames", row.num_frames, unit="frames", per="episode")

    mdr.logger.info(
        "episode={} task={!r} frames={} cameras={}",
        row.episode_id,
        row.task,
        row.num_frames,
        sorted(row.videos),
    )

    return row


pipeline = (
    mdr.read_lerobot("hf://datasets/macrodata/aloha_static_battery_ep005_009")
    .map(log_stats)
)

pipeline.launch_local(name="quickstart-aloha-summary")

Local launch is the standard way to run a complete Refiner pipeline. It distributes shards across worker processes on your machine.

Advanced example

Here is a more elaborate example, reading from and writing to Hugging Face buckets.

import refiner as mdr

dataset = "hf://datasets/yifengzhu-hf/LIBERO-datasets/libero_spatial"
# Replace this with your output bucket (HF, S3, GCP, etc).
output = "hf://buckets/macrodata/test_bucket/libero-spatial"

(
    mdr.read_hdf5(
        dataset,
        groups="/data/demo_*",
        datasets={
            "raw_action": "actions",
            "observation.images.image": "obs/agentview_rgb",
            "observation.images.wrist_image": "obs/eye_in_hand_rgb",
            "ee_state": "obs/ee_states",
            "gripper_state": "obs/gripper_states",
        },
        file_path_column="file_path",
        cache_remote_files=True,
    )
    .map(
        lambda row: row.update(
            task=str(row["file_path"])
            .rsplit("/", 1)[-1]
            .removesuffix(".hdf5")
            .removesuffix("_demo")
            .replace("_", " ")
        )
    )
    .to_robot_rows(
        task_key="task",
        fps=10.0,
        robot_type="libero",
        action_key="raw_action",
        state_key=("ee_state", "gripper_state"),
        video_keys={
            "observation.images.image": "observation.images.image",
            "observation.images.wrist_image": "observation.images.wrist_image",
        },
    )
    .write_lerobot(output, max_video_prepare_in_flight=2)
    .launch_local(
        name="libero-spatial-subset",
        num_workers=2,
    )
)

This example converts the public LIBERO spatial HDF5 subset to LeRobot using local workers. It reads one demo group per row, derives the task label from the filename, turns action/state/image arrays into robotics episodes, encodes the two camera streams as videos, and writes a LeRobot dataset to your output bucket.

The input dataset is public, but writing to your Hugging Face bucket requires an HF_TOKEN in the environment inherited by the local workers:

export HF_TOKEN="your-write-token"

For the full four-suite LIBERO conversion, see Libero HDF5.

Where to go next