VLA Toolkit provides small commands for benchmarking VLA policies, working
with LeRobot datasets, and fine-tuning models. It is a uv workspace whose
tools can be used independently.
uv sync --all-packages
uv run pytestuv run vla-benchmark plan \
--checkpoint allenai/MolmoAct2-LIBERO-LeRobot \
--tasks libero_spatial \
--episodes 50 \
--output-dir artifacts/eval/spatialSee tools/benchmark/README.md.
Inspect any local LeRobot v2/v3 dataset:
uv run vla-dataset inspect /datasets/robot-data \
--require-feature observation.state \
--require-feature actionDataset rewrite, validation, and metadata utilities are independent of the fine-tuning command. See tools/dataset/README.md.
The base command accepts the model, dataset path, output path, and training parameters directly as CLI flags:
WANDB_ENTITY=my-team WANDB_PROJECT=robotics \
uv run vla-finetune ma2 \
--run-name experiment-001 \
--model allenai/MolmoAct2 \
--model-revision e432d85f6e039edca44afb93c262f3084ab72a9c \
--dataset /data/lerobot \
--output-dir /output/experiment-001 \
--updates 1000 \
--dataloader-workers 2 \
--preprocess-mode async \
--preprocess-workers 1 \
--wandbThe command knows nothing about cloud volumes or how its paths were mounted.
Credentials are read from environment variables such as WANDB_API_KEY and
HF_TOKEN. --dry-run prints the resolved process command without executing
it. A real run writes the resolved plan into the output directory as diagnostic
output, not as another required configuration input.
Use --profile for phase timings and rank-selected Perfetto traces, and
--flop-count for lower-bound MFU measurement. See
tools/finetune/README.md.
The Modal launcher accepts Modal-specific Volume, Secret, and resource flags, mounts them at ordinary paths, then runs the same MA2 implementation:
MODAL_PROFILE=my-modal-profile MODAL_ENVIRONMENT=main \
uv run --group modal modal run providers/modal_finetune.py::ma2 \
--run-name experiment-001 \
--dataset-volume my-lerobot-datasets \
--dataset-subpath datasets/robot-data \
--cache-volume my-model-cache \
--output-volume my-training-output \
--gpu 'H100!:8' \
--num-processes 8 \
--updates 1000 \
--dataloader-workers 2 \
--preprocess-mode async \
--preprocess-workers 1 \
--launchAnother provider, such as Nebius, can expose its own storage and compute flags without changing the base MA2 command. See providers/README.md.
- Model-specific trainer subcommands accept direct CLI flags.
- Base training code only sees model identifiers and filesystem paths.
- Providers own provisioning, mounts, durability, telemetry, and secrets.
- Credentials arrive through environment variables.
- Generated plans are diagnostic outputs, never required inputs.