Lecture Notes and Experiments on Autonomous Robots
| Topic | Note | Code |
|---|---|---|
| ROS2 | ROS2 | |
| Navigation | - | |
| Localization | - | |
| Perception | - | |
| Visual Geometry | - | |
| Kinematics | - | |
| Trajectory Planning | - | |
| Diffusion and Flow Matching Applied to Behavior Cloning | - | |
| Vision Language Action | - | |
| World Models | - |
- Introduction to Autonomous Robots
- Introduction to Autonomous Mobile Robots
- Reinforcement Leanring: An Overview by Kevin Murphy
- SLAM Handbook
- MIT 16.485 — Visual Navigation for Autonomous Vehicles (VNAV) — mathematical foundations of visual navigation (geometry to optimization), state-of-the-art algorithms, and software packages
- Robot Learning by ETH
- LIBERO
- University of Texas-Dallas
- Flow Matching and Diffusion Models ⭐
- Learning Affordances at Inference-Time for Vision-Language-Action Models
- RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
- Sakana AI Scientist
- AutoResearch
- AutoResearchClaw
- Sakana RSI Lab
- Hierarchical Reasoning Model
- First Steps Toward Automated AI Research
- EgoDex
- Princeton Digital Assets
- GraspGen
- ShareGPT-4o-Image
- Wan Video Generation
- RoboGen
- Large Scale Grasp Dataset
- Rosario Agricultural Dataset
- Project Go-Big
- Behavior Challenge by Stanford
- BridgeData V2: A Dataset for Robot Learning at Scale
- Robocasa 🌟
- Egocentric Dataset
- ACRONYM: A Large-Scale Grasp Dataset Based on Simulation
- Spatial AI
- In-N-On Scaling Egocentric manipulation with in-the-wild and on-task data
- Articraft-10K: Agentic System for Scalable Articulated 3D Asset Generation ⭐
- HumanNet: Scaling Human-Centric Video Learning to One Million Hours ⭐
- GR00T Dataset ⭐
- Scalable Behavior Cloning with Open Data, Training, and Evaluation
- LightwheelAI EgoDemo ⭐ — 50-hour developer sample of EgoSuite-Open100K (100K hours, 15,000+ tasks/scenes), the largest fully-annotated open egocentric human dataset: hand + full-body pose, semantic annotations, in LeRobot v3 / MCAP / raw MP4 (gated access)
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
- Diffusion vs Autoregressive in Robotics
VLA evaluation is moving beyond LIBERO success-rate leaderboards toward a broader stack: unified sim-benchmark harnesses, real-robot production metrics, and independent real-world evaluation.
- How to Evaluate General-Purpose Robot Policies for Real-World Deployment (NVIDIA)
- VLA Evaluation Harness (Allen AI) — one framework to evaluate any VLA model on any robot simulation benchmark: 18+ benchmarks (LIBERO, SimplerEnv, CALVIN, ManiSkill2, RoboCasa, RoboTwin, RLBench, ...) behind a single interface, with a VLA leaderboard.
- LeRobot Evaluation —
lerobot-evalgives one evaluation interface across multiple sim benchmarks, each wrapped as a Gymnasium environment behind a standardgym.Envinterface. - PhAIL — Physical AI Leaderboard — real-robot benchmark (Franka FR3) scored on production metrics like throughput and failures; distributional methodology (time-to-success CDF, Human-Relative Throughput, KS significance tests). Paper
- Robocurve — independent, real-world robot evaluation: open-source Inspect Robots framework (any model × any embodiment × any benchmark, with full trace logs and Rerun visualization) plus the World Evals benchmark catalog.
- RoboDojo — brings sim + real-world evaluation together: 42 sim tasks and 18 real-world tasks across 3 embodiments, five capability dimensions, heterogeneous parallel simulation in Isaac Sim, and a reproducible RealEval system. Paper
- FoundationGrasp and GraspGPT
- Visual Dexterity: In-Hand Reorientation of Novel and Complex Object Shapes
- AffordGrasp
- Grasp Diffusion
- Open-Vocabulary Affordance Detection in 3D Point Clouds
- GraspGen
- OpenArm - Open Source Humanoid
- All Humanoids as of Oct 2025
- Awesome Humanoid Learning
- WUJI Hand as used by Genesis AI
- Isaac GROOT
- Noetix Robotics
- Humanoid Robots - Not Solution
- RoboMME: A Large-Scale Standardized Benchmark for VLA Memory ⭐
- Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering
- NVIDIA's ReMEmbR
- FindingDory
- VLAs Memory
- Observational Memory
- EasyR1
- R-Zero
- verl: Volcano Engine Reinforcement Learning for LLMs
- UWLab and Omni-Reset
- Reinforcement Learning Book by K Murphy
- Agentic RL
- NVIDIA Molt Agent RL Framework ⭐
- The State of Simulation for Physical AI: An Overview (NVIDIA)
- Habitat
- Newton
- Viral Humanoid
- Molmo-Bot
- Real2Sim2Real-SimFoundry
- Genesis World (supports touch/tactile
- Spatially-anchored Tactile Awareness for Robust Dexterous Manipulation
- Intrinsic sense of touch for intuitive physical human-robot interaction
Detailed summaries and resources can be found in the World Models Documentation
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World Action Models
- World Action Models: A Survey — Shen et al., 2026 (57 pages). A comprehensive survey clarifying the boundaries among world models, video generation models, action-grounded video world models, VLA policies, and World Action Models (WAMs). WAMs are embodied predictive-action models that forecast the future to inform action. The survey organizes existing work through two views: (1) what each method generates — rendered futures, latent futures, or video-generation-free action reasoning; and (2) architectural decomposition by predictive substrate, backbone, action coupling, and deployment regime. Key themes include interactability, causality, persistence, physical plausibility, and generalization. The emerging design pattern: WAMs are generating less of the future while preserving what control requires, trading representational richness against compute, memory, latency, and action-label cost. Homepage