You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Private real-robot deployment bundle. Access is granted manually by the owner.
Log in or Sign Up to review the conditions and access this model content.
Long-WAM on the Unitree G1
Same two-process layout as the pi0.5 deployment: the policy runs on the workstation GPU (RTX 5090) behind a WebSocket, the robot PC streams observations and applies joint targets.
robot PC (eval_longwam_g1.py) --30 Hz obs (3 JPEG + 16-D state)--> workstation (policy_server.py)
<-- 32 x 16 joint targets every replan -- LongWAMPolicy + accel
Files
| file | runs on | purpose |
|---|---|---|
policy_server.py |
workstation | loads a checkpoint, keeps the 48-step observation history, serves WebSocket |
longwam_policy.py |
workstation | model wrapper: pinzi frame composition (identical to training), normalization, infer_joint_ar, acceleration switches |
bench_policy.py |
workstation | latency / memory / parity benchmark of the acceleration modes |
eval_longwam_g1.py |
robot PC | pi0.5-style control loop (unitree_lerobot cameras + arms + Dex1) |
contract.py |
both | the fixed I/O contract (16-D order, camera mapping, 384x320 layout, 48/32 timing) |
msgpack_numpy.py, websocket_client.py |
both | OpenPI-compatible wire protocol, no torch on the robot |
make_deploy_bundle.sh |
cluster | packs code + per-task model files into one directory to copy |
bench_g1_policy.sbatch, loopback_g1_policy.sbatch |
cluster | the tests that were run on H100 |
Contract (must match training)
- Cameras:
observation/color_0= head left eye (top tile),observation/color_2= left wrist,observation/color_3= right wrist; RGB uint8 480x640, raw array or JPEG bytes. - State / action: 16 values
[left_arm q0..q6, right_arm q0..q6, left_dex1, right_dex1], absolute targets in native units; a chunk is 32 steps at 30 Hz. - The server owns the history: send an observation every control step
(
mode: "observe"), ask for a chunk withmode: "infer", sendreset: trueat episode start. - Per task you need:
config.yamlof the training run, onestep_XXXXXX.pt, the task'sfastwam_dataset_stats.json, and the text cache directory (1 MB; avoids the 11 GB T5). - Tasks:
Dynamic_Cup(both arms),Speed_Cup_20/30/40/50(right arm; conveyor speed in the name). All share the promptpick up cube.; pick the checkpoint by task.
1. Build the bundle (cluster)
bash FastWAM/deploy/unitree_g1/make_deploy_bundle.sh /path/to/longwam_g1_bundle \
Dynamic_Cup step_020000.pt Speed_Cup_20 step_020000.pt # any list of <task> <weights>
The bundle also contains the Wan2.2 VAE under models/Wan-AI/Wan2.2-TI2V-5B/ (2.8 GB). Copy the whole
bundle to the workstation and bundle/robot/ to the robot PC. The same layout is published at
https://huggingface.co/AaronHuangWei/Long-WAM-G1-Dynamic-Task-Deploy (gated; request access).
2. Workstation (RTX 5090)
Requirements: Python 3.10/3.11, PyTorch 2.7.1+cu128 (the training runtime), then
pip install -e code/FastWAM websockets msgpack (Pillow decodes JPEG; it comes with torchvision).
cd bundle/code/FastWAM
export DIFFSYNTH_MODEL_BASE_PATH=$PWD/../../models # contains Wan-AI/Wan2.2-TI2V-5B/Wan2.2_VAE.pth
export PYTHONPATH=src:deploy/unitree_g1
export TORCHINDUCTOR_CACHE_DIR=$HOME/.cache/longwam_inductor # compiled kernels survive restarts
python deploy/unitree_g1/policy_server.py \
--run-dir ../../models/Dynamic_Cup \
--checkpoint ../../models/Dynamic_Cup/step_020000.pt \
--stats ../../models/Dynamic_Cup/fastwam_dataset_stats.json \
--text-cache-dir ../../text_embeds_cache \
--accel lossless --port 8000
The server prints serving Long-WAM G1 policy at ws://... after warm-up. First start with a
cold Inductor cache compiles for several minutes; afterwards start-up is about one minute.
--accel:
| value | what | numerics | H100 p50 (step-2877 ckpt) |
|---|---|---|---|
eager |
plain BF16 | reference | 478 ms |
lossless (default) |
resident RoPE + torch.compile of VAE encode, video denoiser, video KV prefill, action denoiser; --compile-mode max-autotune = CUDA Graphs |
max abs action diff vs eager 0.024 (BF16 noise) | 173 ms |
nvfp4 |
lossless + NVFP4 (W4A4) video expert via FourOverSix; action expert stays BF16, KV cache stays BF16 |
lossy, opt-in | needs sm_120 (5090 only) |
Use --compile-mode max-autotune-no-cudagraphs to keep the compile speed-up without CUDA Graphs.
--no-gpu-check allows GPUs other than the 5090 (compute capability 12.0).
NVFP4 additionally needs FourOverSix built for sm_120 with the three patches in
third_party/patches/; run scripts/build_nvfp4_sm120.sh on the workstation
(see docs/nvfp4_environment.md).
Verify before driving the robot (prints p50/p95, peak memory, max-abs action difference vs eager):
python deploy/unitree_g1/bench_policy.py --run-dir ../../models/Dynamic_Cup \
--checkpoint ../../models/Dynamic_Cup/step_020000.pt \
--stats ../../models/Dynamic_Cup/fastwam_dataset_stats.json \
--text-cache-dir ../../text_embeds_cache --accel eager lossless nvfp4
3. Robot PC
Requirements: the unitree_lerobot environment (its eval_robot/make_robot.py) plus
websockets msgpack numpy opencv-python. No torch.
cd unitree_lerobot
PYTHONPATH=/path/to/bundle/robot:$PYTHONPATH python /path/to/bundle/robot/eval_longwam_g1.py \
--policy_server_host WORKSTATION_IP --policy_server_port 8000 \
--arm G1_29 --ee dex1 --control_side both --replan_steps 8 \
--send_real_robot true --motion true
--control_side right|left|both: which arm(s) follow the policy; the other holds its measured pose (Speed_Cup data only moves the right arm).--replan_steps 8: execute 8 of the 32 predicted steps, then request a new chunk (every 0.27 s; keep it above server latency x 30 Hz).--image_format jpeg(default, ~180 KB/step) orraw(2.7 MB/step).- Safety: a per-step joint jump above
--max_arm_target_delta(rad) or--max_gripper_target_deltaaborts the loop, as in the pi0.5 script. - The Dex1 access (
ee_shared_mem["state"]=[left, right],ee_shared_mem["left"/"right"]targets of length 1) follows unitree_lerobot's gripper interface; check it against the installedmake_robot.pybefore the first run.
What was tested
- CPU: contract, protocol pack/unpack, JPEG decode, pinzi composition, config building (
SMOKE OK); the ported acceleration unit tests (tests/test_nvfp4_inference_optimization.py,tests/test_rope_real_pair.py) and the pre-existing model tests pass. - H100 (
bench_g1_policy.sbatch): eager vs lossless numbers above. - H100 (
loopback_g1_policy.sbatch):policy_server.py(--accel lossless) driven by a synthetic JPEG client through the real WebSocket path: 64 steps, 8 chunks, client round-trip p50 184 ms of which server inference 172 ms. - Not tested here: NVFP4 (requires the 5090), the real Unitree hardware loop.