Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VLA Toolkit

VLA Toolkit provides small commands for benchmarking VLA policies, working with LeRobot datasets, and fine-tuning models. It is a uv workspace whose tools can be used independently.

uv sync --all-packages
uv run pytest

Benchmark

uv run vla-benchmark plan \
  --checkpoint allenai/MolmoAct2-LIBERO-LeRobot \
  --tasks libero_spatial \
  --episodes 50 \
  --output-dir artifacts/eval/spatial

See tools/benchmark/README.md.

Dataset

Inspect any local LeRobot v2/v3 dataset:

uv run vla-dataset inspect /datasets/robot-data \
  --require-feature observation.state \
  --require-feature action

Dataset rewrite, validation, and metadata utilities are independent of the fine-tuning command. See tools/dataset/README.md.

Fine-tune

The base command accepts the model, dataset path, output path, and training parameters directly as CLI flags:

WANDB_ENTITY=my-team WANDB_PROJECT=robotics \
uv run vla-finetune ma2 \
  --run-name experiment-001 \
  --model allenai/MolmoAct2 \
  --model-revision e432d85f6e039edca44afb93c262f3084ab72a9c \
  --dataset /data/lerobot \
  --output-dir /output/experiment-001 \
  --updates 1000 \
  --dataloader-workers 2 \
  --preprocess-mode async \
  --preprocess-workers 1 \
  --wandb

The command knows nothing about cloud volumes or how its paths were mounted. Credentials are read from environment variables such as WANDB_API_KEY and HF_TOKEN. --dry-run prints the resolved process command without executing it. A real run writes the resolved plan into the output directory as diagnostic output, not as another required configuration input.

Use --profile for phase timings and rank-selected Perfetto traces, and --flop-count for lower-bound MFU measurement. See tools/finetune/README.md.

Modal

The Modal launcher accepts Modal-specific Volume, Secret, and resource flags, mounts them at ordinary paths, then runs the same MA2 implementation:

MODAL_PROFILE=my-modal-profile MODAL_ENVIRONMENT=main \
uv run --group modal modal run providers/modal_finetune.py::ma2 \
  --run-name experiment-001 \
  --dataset-volume my-lerobot-datasets \
  --dataset-subpath datasets/robot-data \
  --cache-volume my-model-cache \
  --output-volume my-training-output \
  --gpu 'H100!:8' \
  --num-processes 8 \
  --updates 1000 \
  --dataloader-workers 2 \
  --preprocess-mode async \
  --preprocess-workers 1 \
  --launch

Another provider, such as Nebius, can expose its own storage and compute flags without changing the base MA2 command. See providers/README.md.

Design rules

  1. Model-specific trainer subcommands accept direct CLI flags.
  2. Base training code only sees model identifiers and filesystem paths.
  3. Providers own provisioning, mounts, durability, telemetry, and secrets.
  4. Credentials arrive through environment variables.
  5. Generated plans are diagnostic outputs, never required inputs.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages