Skip to content

Repository files navigation

Physical AI Operating System

We're building the "Android" for physical AI agents.

Give Hermes, Claude Code, Codex, and other AI agents eyes, ears, a voice, and a body they can control.

Bring your AI agents into the physical world.

Start with Lamp · Bring your own robot · Build a physical skill

Autonomous.Lamp.mp4

Choose your agent

Run Hermes, Claude Code, Codex, OpenClaw, OpenCode, or PicoClaw. Choose a compatible model and configure its voice, personality, and tools.

The OS runs on the robot and coordinates hardware and agent tasks. Model inference may use remote services, depending on your configuration.

Give it a body

Start with Lamp, Intern, or Reachy Mini—or add support for your own robot. Physical AI Operating System connects the agent to the hardware that body provides: cameras, microphones, speakers, motors, lights, and sensors.

Physical skills

Give your agent new ways to sense and interact with its surroundings.

  • Follow an object with its camera and movement.
  • Respond to touch with a gesture.
  • Turn a sensor reading into light, movement, or speech.

Start with an existing physical skill, change its instructions, or connect your own tools. Use the skill creator to write a new skill, and see the contribution guide to share it.

Try saying

On a configured Lamp, try these in normal voice mode or text-only Web/MQTT chat. These are supported intent examples; selection depends on confidence and available device capabilities. Uncertain requests go to the main agent.

Say or type Expected behavior
“Turn off the lights.” Turn off Lamp's light through a local command.
“This lamp is too bright.” Reduce the current light brightness by half.
“You're speaking too loudly.” Reduce the current speaker volume by half.
“Make this lamp violet.” Set a solid purple light.
“I need light to read a book.” Activate the reading lighting scene.

Context matters too: while working on an image or render, “Make it brighter” goes to the main agent to interpret the task context. It does not automatically brighten the Lamp. Explicit hardware requests such as “Turn off the lights” remain eligible for intent handling while a Harness task is pending.

Intent configuration and limits · Harness context routing

Work across your robot and computer

  • Delegate to Harness agents. harness-use sends coding and research tasks to agents already running in Harness on your paired computer. Ask a named agent to work, answer its follow-up questions, and receive its result through voice or chat. Pair from OS Monitor and Harness Desktop on the same LAN. Harness-only voice sends manual tap-to-record turns to the agent focused in Harness. Get Harness · Integration and setup.
  • Use Mac apps through Buddy. computer-use lets the device agent inspect an app, click or type, and verify the resulting UI through a paired Autonomous Buddy. Buddy bundles Cua Driver for Accessibility observations and actions, with screenshot support for visual tasks. The device agent owns the task; Buddy executes on the Mac. Harness and Buddy have separate connections and pairing. Computer-use guide.
  • Route requests with Jev. For normal voice requests received by os-server and text-only Web/MQTT chat, local rules run first, then the Jev intent fallback (enabled by default). The OS validates the selected command, parameters, confidence, and device capabilities before execution; uncertain requests continue to the main agent. Context-dependent follow-ups bypass intent handling when they need the main agent. Attachments and Harness-only voice retain their separate routes. Separate skill-preloading integrations prepare instructions before the main model call through runtime hooks or managed bridges, with native discovery as fallback. New integrations outside Hermes are disabled by default pending native validation, controlled by a Go build switch in each runtime. Jev also supports experimental Buddy UI-action suggestions.

Quick start

The simplest way in is a robot we have already tested it on. What each of them can do: robot comparison.

Autonomous Lamp

Lamp is the robot that shows the whole OS — it sees, hears, speaks, moves, and ships with Physical AI Operating System on it.

  1. Add it. In the Autonomous app (iOS | Android), tap Add robot → Lamp.
  2. Set up Wi-Fi. Pick your network in the app; it joins the robot's hotspot and hands over the keys and pairing.
  3. Interact with Lamp. Say something, it turns to look at you, the ring lights up, and it answers.
  4. Install a skill from the Skill Store — one tap, live on the next conversation.
  5. Build your own skill. Type what you want it to do in the app and it writes the skill.
  6. Give it a character. Edit SOUL.md and it is someone else on the next turn.

Reachy Mini

Reachy Mini is Hugging Face's desk robot, running our OS beside its own stack.

Reach.Mini.and.Autonomous.Lamp.mp4
  1. SSH in — ssh pollen@reachy-mini.local.
  2. Run one command. Nothing is flashed; the Reachy daemon keeps the motors.
    curl -fsSL https://raw.githubusercontent.com/autonomous-ai/Physical-AI-Operating-System/main/robots/reachy-mini/install.sh | sudo bash
  3. Add it. In the app, tap Add robot → Reachy Mini and give it reachy-mini.local.
  4. Interact with it. Say something — the head tilts, the antennas lift, and it answers.
  5. Install a skill from the Skill Store, or type what you want it to do and it writes one.
  6. Give it a character. Edit /opt/devices/reachy-mini/SOUL.md. Everything else, including how to undo the install: devices/reachy-mini/README.md.
  7. Put it next to a Lamp. Each one hears the other's answer as its next input, so the two of them will hold a conversation until you stop them.

Autonomous Intern

Intern is the always-on desk agent: mic, speaker, LED ring.

Autonomous Intern on a desk beside a laptop, tip glowing blue

  1. Add it. In the app, tap Add robot → Intern.
  2. Set up Wi-Fi. Same flow as Lamp: pick your network and it handles the keys and pairing.
  3. Interact with it. Say something and it answers; the ring shows what it is doing.
  4. Install a skill from the Skill Store.
  5. Build your own skill. Type what you want in the app; it is live on the next conversation.
  6. Give it a character. Edit /opt/devices/intern-v2/SOUL.md.

Bring your own robot

Physical AI Operating System runs on any robot you can describe in four markdown files.

  • ROBOT.md — the body: the board and the hardware it has.
  • SOUL.md — the self: who it is and how it talks.
  • SAFETY.md — the bounds: how fast, how bright, how late.
  • SKILL.md — the hands: one thing it can do.

Follow the full guide.

Platform architecture

One voice, three ways to act

A conversation can stay with realtime. A music request can use music and audio skills to play a song at the requested volume, with LED feedback on supported bodies. A computer task can use Harness to reach a paired agent, then return a short spoken result.

One shared Physical AI Operating System diagram with three routes: realtime answers conversation directly; the main runtime uses music and audio skills to play jazz at 30% volume with HAL LED feedback; or it uses harness-use to update a Blender scene, whose final result returns through OS to HAL for speech.

Download the SVG and open it in a browser to see the animation. GitHub shows the static diagram; timing is illustrative.

The layers behind it

Physical AI Operating System is a software stack. Each layer uses only the layer below it, so any layer can be replaced without touching the others. Every layer is a folder in this repo.

Physical AI Operating System stack, top down: apps, skills, the agentic runtime, the Go system services, the realtime voice agent, the capabilities a robot declares, the safety gate, drivers, boards, the vendor Linux kernel, and the bodies — one colour per layer, and the rows you can extend yourself drawn dashed

View the diagram. Layers show the platform structure, not request execution order.

What a person touches. The Autonomous app adds a robot, sets up Wi-Fi, installs skills from the Skill Store and switches brains; the robot also serves its own setup and monitor UI from system/web/. Both talk to os-server on :5000.

One folder per behavior, one SKILL.md inside: markdown the agent reads. Hardware skills emit [HW:/path:{json}] markers for OS dispatch instead of touching a servo bus or GPIO pin. Computer and agent skills use helpers that call OS APIs and return observations or task results. Skills declare required capabilities so they install on compatible robots.

The engine that thinks. Six of them — Hermes, OpenClaw, PicoClaw, Codex, Claude Code, OpenCode — behind one 76-method AgentGateway. It reads the robot's SOUL.md and its installed skills. Switch live from the web UI; persona, memory and connectors move with it.

The Go daemon os-server on :5000, one package per box in the figure. intent handles eligible voice and text-only Web/MQTT commands through local rules, then a validated Jev fallback (on by default), while deferring contextual follow-ups to the main agent; harness delegates work to paired computer agents; buddy carries Mac observations and actions; server strips [HW:…] markers out of a reply and POSTs them to HAL before the words are spoken; agent switches engines; bootstrap is OTA, its own binary.

HAL supports Gemini Live (default model: gemini-3.8-live), OpenAI Realtime, GPT-Live (gpt-live-1), and Pipecat v1 (pipecat_v1). In normal voice mode, the realtime agent answers directly or delegates tasks to the main runtime; Harness-only voice routes manual captures directly to Harness.

Pipecat runs the pipeline inside HAL: speech recognition → an OpenAI-compatible LLM (default: qwen/qwen3.6-35b-a3b) → text spoken by HAL's TTS. Pipeline orchestration runs on the robot; model calls still use remote services. It supports both committed turns and continuous Live input.

Smart Turn runs a local ONNX model to help decide when a person has finished speaking. It supplements silence detection in the shared non-Live capture path and works with Silero VAD in Pipecat Live. Bounded silence fallbacks handle unavailable inference; manual Harness recording still ends on the user's tap. Pipecat and Smart Turn require HAL's optional pipecat extra, included in Lamp/Pi/OrangePi setup but excluded from Reachy because of conflicting ONNX dependencies. Voice architecture, configuration and limits.

The 13 names a robot may declare — audio, vision, sensing, presence, motion, light, display, expression, lifelike, media, connectivity, companion, system. Ten mount HTTP routes on :5001 (111 endpoints, live Swagger at /api/hardware/docs); presence and lifelike are loops with no route, companion lives in os-server. HAL mounts only what ROBOT.md declares and fails loud on a missing required driver.

A pure function of SAFETY.md, below the engine and in every request path: brightness, quiet hours, explicit-move speed. No model in the loop — the same clamp whoever asked. What it does not cover yet: docs/safety.md.

One folder per subsystem: motors, rgb, camera, voice, display, sensing, tracking, and the media handover a third-party daemon needs. New hardware is one class and one factory line.

One JSON entry per board, matched against /proc/device-tree/model. Raspberry Pi 4, Pi 5, CM4 and OrangePi 4 Pro today. A new board is an entry, not a code change.

Linux

The vendor kernel — Raspberry Pi OS, OrangePi Debian, or the robot's own image. We do not ship one, and nothing above the drivers has a real-time deadline: position control closes in the servo firmware, or in the robot's own daemon.

Four markdown files and a driver per robot. Declarations, not forks — a body is a PR.

Long form: architecture · HAL · device spec · capabilities · safety · developer guide.

Contribute

The easiest way in is a skill: one markdown file, no Go, no hardware, and it lands on every robot that has the parts. PRs welcome, vibe-coded ones included. Questions, half-built ports and show-and-tell go in Discussions; gaps we would love help with are labelled claim-me — comment to take one.

You want to… You write… Start from
Teach every robot something new skills/<name>/SKILL.md (+ skill.json if it needs hardware) skills/guard/ · skill-creator
Run the OS on your robot robots/<id>/ROBOT.md + SAFETY.md + SOUL.md robots/reachy-mini/ — a third-party port, end to end
Support new hardware a class in hal/drivers/<subsystem>/ + one factory line reachy_service.py
Support a new board one entry in hal/board/boards.json boards.json
Add a brain an AgentGateway implementation in runtimes/<name>/ adding-agent-runtime.md

Seven more paths — apps, chat bridges, perception models, voices, safety bounds, CTS probes — and the norms: CONTRIBUTING.md. One rule worth knowing up front: robots/contract/ is the interface everyone builds on, so open an issue before you change it.

Build locally:

make os-build && make os-test          # Go daemon, cross-compiled to linux/arm64
(cd hal && uv sync) && make hal-dev    # HAL on :5001 with reload
make web-install && make web-dev       # setup + monitor UI
make cts                               # is this a valid Autonomous device?

License

Everything outside hal/ is Apache-2.0. hal/ is GPL-3.0, kept that way by choice so the tree has one license per top-level folder; a driver you commit there is GPL, so a closed vendor SDK wraps out of process.

A robot running this carries other people's work: Pollen's reachy_mini SDK, YOLOv8 for tracking (AGPL-3.0 — read it before you ship), TEN-VAD and Silero for hearing, LeRobot and the LeLamp Runtime under the motion code, and the brains we install but do not ship. All of it, including what we copied verbatim: CREDITS.md. Security issues: SECURITY.md.