π₀.₅ is a
Vision-Language-Action model with open-world generalization
, from Physical Intelligence. The LeRobot implementation is adapted from their open source
OpenPI
repository.
Model Overview
π₀.₅ represents a significant evolution from π₀, developed by
Physical Intelligence
to address a big challenge in robotics:
open-world generalization
. While robots can perform impressive tasks in controlled environments, π₀.₅ is designed to generalize to entirely new environments and situations that were never seen during training.
The Generalization Challenge
As Physical Intelligence explains, the fundamental challenge isn't performing tasks of agility or dexterity, but generalization, the ability to correctly perform tasks in new settings with new objects. Consider a robot cleaning different homes: each home has different objects in different places. Generalization must occur at multiple levels:
Physical Level
: Understanding how to pick up a spoon (by the handle) or plate (by the edge), even with unseen objects in cluttered environments
Semantic Level
: Understanding task semantics, where to put clothes and shoes (laundry hamper, not on the bed), and what tools are appropriate for cleaning spills
Environmental Level
: Adapting to "messy" real-world environments like homes, grocery stores, offices, and hospitals
Co-Training on Heterogeneous Data
The breakthrough innovation in π₀.₅ is
co-training on heterogeneous data sources
. The model learns from:
Multimodal Web Data
: Image captioning, visual question answering, object detection
Verbal Instructions
: Humans coaching robots through complex tasks step-by-step
Subtask Commands
: High-level semantic behavior labels (e.g., "pick up the pillow" for an unmade bed)
Cross-Embodiment Robot Data
: Data from various robot platforms with different capabilities
Multi-Environment Data
: Static robots deployed across many different homes
Mobile Manipulation Data
: ~400 hours of mobile robot demonstrations
This diverse training mixture creates a "curriculum" that enables generalization across physical, visual, and semantic levels simultaneously.
Training
Here's a complete training command for finetuning the base π₀.₅ model on your own dataset:
pi05_base huggingface.co is an AI model on huggingface.co that provides pi05_base's model effect (), which can be used instantly with this lerobot pi05_base model. huggingface.co supports a free trial of the pi05_base model, and also provides paid use of the pi05_base. Support call pi05_base model through api, including Node.js, Python, http.
pi05_base huggingface.co is an online trial and call api platform, which integrates pi05_base's modeling effects, including api services, and provides a free online trial of pi05_base, you can try pi05_base online for free by clicking the link below.
lerobot pi05_base online free url in huggingface.co:
pi05_base is an open source model from GitHub that offers a free installation service, and any user can find pi05_base on GitHub to install. At the same time, huggingface.co provides the effect of pi05_base install, users can directly use pi05_base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.