LeRobot documentation
EO-1
EO-1
EO-1 is a Vision-Language-Action policy for robot control. The LeRobot implementation integrates EO-1 with the standard LeRobot training, evaluation, processor interface.
Model Overview
EO-1 uses a Qwen2.5-VL backbone for vision-language understanding and adds a continuous flow-matching action head for robot control. The policy formats each robot-control sample as a multimodal conversation: camera images are passed to Qwen2.5-VL, the robot state is represented with EO-1 state tokens, and the future action chunk is represented with EO-1 action tokens.
During training, EO-1 learns to denoise continuous action chunks at the action-token positions. During inference, it samples an action chunk, returns continuous actions, and executes n_action_steps from the chunk before sampling again.
What the LeRobot Integration Covers
- Standard
policy.type=eo1configuration through LeRobot - Qwen2.5-VL image and text preprocessing through policy processors
- Continuous flow-matching action prediction
- Checkpoint save/load through LeRobot policy APIs
- Training with
lerobot-trainand evaluation withlerobot-eval
The broader EO-1 project also includes interleaved vision-text-action pretraining and multimodal reasoning workflows. This page focuses on the LeRobot robot-control policy path.
Installation Requirements
Install LeRobot by following the Installation Guide.
Install EO-1 dependencies by running:
pip install -e ".[eo1]"If you want to train or evaluate on LIBERO, install the LIBERO dependencies too:
pip install -e ".[eo1,libero]"
EO-1 can use the standard PyTorch scaled-dot-product attention backend through policy.attn_implementation=sdpa. If your environment has a compatible flash_attn installation, you can request policy.attn_implementation=flash_attention_2.
Data Requirements
EO-1 expects a LeRobot dataset with:
- At least one visual observation, for example
observation.images.image observation.stateaction- A language task instruction through the dataset
taskfield
If your dataset uses different observation names, use rename_map to align them with the names expected by your training or evaluation setup.
Usage
To use EO-1 in a LeRobot configuration, specify the policy type as:
policy.type=eo1Use the converted release checkpoint to start from EO-1’s interleaved vision-text-action pretraining:
--policy.path=lerobot/eo1-base
lerobot/eo1-base is a native LeRobot conversion of IPEC-COMMUNITY/EO-1-3B.
It preserves the released Qwen vision-language weights and the trained
flow-matching action head. Its action chunk size is 16, matching the upstream
checkpoint. To initialize only from the original Qwen backbone instead, omit policy.path and set policy.vlm_base=Qwen/Qwen2.5-VL-3B-Instruct.
Load another fine-tuned LeRobot-format EO-1 checkpoint with:
policy.path=your-org/your-eo1-checkpoint
Text generation and interactive control
EO-1 exposes the Qwen backbone through the shared policy text contract. Runtime
adds query_kind and text, then runs the normal EO-1 input processor:
request = {
**observation,
"query_kind": "vqa",
"text": "Which object is closest to the gripper?",
}
processed_batch = preprocessor(request)
answer = policy.generate_text(processed_batch)vqa preserves the caller’s question. next_subtask applies the recipe stored
in the EO-1 policy config. EO-1 inherits generate_text(processed_batch) and
overrides supports_text_generation and generate_text; no separate VLM is loaded.
With the interactive language runtime, EO-1 can generate a subtask and rebuild its action prompt from that current subtask:
lerobot-rollout \
--policy.path=your-org/your-eo1-checkpoint \
--policy.device=cuda \
--task="clear the table" \
--interactive=trueAfter /start, use /vqa <question>, /autosteer <goal>, or /subtask <text>.
Training
Training Command Example
lerobot-train \
--dataset.repo_id=your_org/your_dataset \
--policy.path=lerobot/eo1-base \
--policy.dtype=bfloat16 \
--policy.attn_implementation=sdpa \
--policy.gradient_checkpointing=false \
--output_dir=./outputs/eo1_training \
--job_name=eo1_training \
--steps=300000 \
--batch_size=16 \
--policy.device=cudaTrain with language annotations
EO1Config includes a recipe to train from LeRobot’s optional
language columns. The built-in joint recipe supervises a selected assistant
response and its continuous action chunk in the same sequence:
lerobot-train \ --dataset.repo_id=your_org/your_language_annotated_dataset \ --policy.type=eo1 \ --policy.text_loss_weight=0.01 \ --policy.flow_loss_weight=1.0 \ --policy.device=cuda \ --batch_size=1 \ --steps=20000
Targeted assistant turns receive next-token cross-entropy, while low_level recipe streams receive flow-action supervision. A row may receive either
objective or both. User prompts, images, state tokens, action tokens, and
padding are excluded from text loss. Set policy.recipe=null to disable recipe rendering and use the original
action-only training path.
See Language columns and recipes for the annotation
schema and recipe structure. recipe_path may point to an explicit external
YAML override; LeRobot does not provide shared default recipe files.
The default system prompt is part of EO1’s recipe and can be replaced by a custom
system turn. Rendered inputs without a system turn, including VQA prompts,
remain without one. The action-only path without rendered messages keeps the
default system prompt. EO1PrepareModelMessagesStep adds images and state/action tokens
to the rendered messages before Qwen formatting and tokenization.
Key Training Parameters
| Parameter | Default | Description |
|---|---|---|
policy.path | lerobot/eo1-base | Released EO-1 weights converted to native LeRobot format |
policy.vlm_base | Qwen/Qwen2.5-VL-3B-Instruct | Qwen processor/backbone used only when training without policy.path |
policy.dtype | auto | Backbone dtype request: auto, bfloat16, or float32 |
policy.attn_implementation | None | Optional Qwen attention backend, such as sdpa |
policy.gradient_checkpointing | false | Reduces memory usage during training |
policy.chunk_size | 8 | Number of future actions predicted per chunk |
policy.n_action_steps | 8 | Number of actions consumed from a sampled chunk |
policy.num_denoise_steps | 10 | Number of flow-matching denoising steps used during sampling |
policy.max_state_dim | 32 | State padding dimension |
policy.max_action_dim | 32 | Action padding dimension |
policy.force_fp32_autocast | true | Keeps the flow head in fp32 even when the backbone uses mixed precision |
policy.supervise_padding_action_dims | true | Controls whether padded action dimensions are supervised |
policy.supervise_padding_actions | true | Controls whether padded future action rows are supervised |
policy.recipe_path | None | Enables recipe-driven text and per-sample action supervision |
policy.text_loss_weight | 0.01 | Weight for supervised assistant-token cross-entropy |
policy.flow_loss_weight | 1.0 | Weight for continuous flow-action loss |
policy.tokenizer_max_length | 1000 | Maximum joint vision-language-action sequence length |
Evaluation
EO-1 can be evaluated through lerobot-eval once you have a LeRobot-format checkpoint:
lerobot-eval \ --policy.path=your-org/your-eo1-checkpoint \ --env.type=libero \ --env.task=libero_object \ --eval.batch_size=1 \ --eval.n_episodes=20
For datasets or environments whose camera names differ from the checkpoint configuration, pass a rename_map:
lerobot-eval \
--policy.path=your-org/your-eo1-checkpoint \
--env.type=libero \
--env.task=libero_object \
--rename_map='{"observation.images.image2":"observation.images.wrist_image"}'Configuration Notes
Image Processing
EO-1 uses the Qwen2.5-VL processor. The policy.image_min_pixels and policy.image_max_pixels settings control the image resizing bounds before the visual tokens are passed into the backbone.
State and Action Dimensions
The policy pads state and action vectors to policy.max_state_dim and policy.max_action_dim before the EO-1 flow head. Predictions are cropped back to the original action dimension before being returned by the policy.
Attention Backend
Use policy.attn_implementation=sdpa for a portable setup. Use flash_attention_2 only when flash_attn is installed and compatible with your environment.
References
Citation
@article{eo1,
title={EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control},
author={Delin Qu and Haoming Song and Qizhi Chen and Zhaoqing Chen and Xianqiang Gao and Xinyi Ye and Qi Lv and Modi Shi and Guanghui Ren and Cheng Ruan and Maoqing Yao and Haoran Yang and Jiacheng Bao and Bin Zhao and Dong Wang},
journal={arXiv preprint},
year={2025},
url={https://arxiv.org/abs/2508.21112}
}License
This LeRobot integration follows the Apache 2.0 License used by LeRobot. Check the upstream EO-1 model and dataset pages for the licenses of released EO-1 checkpoints and data.
Update on GitHub