4DAnyone: Create Anyone in 4D from a Casual Monocular Video Paper • 2608.20335 • Published 18 days ago • 83
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization Paper • 2608.26103 • Published 12 days ago • 26
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
From Foundation to Application: Improving VLA Models in Practice Paper • 2607.06403 • Published Jul 7 • 20
Geometric Context Transformer for Streaming 3D Reconstruction Paper • 2604.14141 • Published Apr 15 • 38
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Paper • 2604.20796 • Published Apr 22 • 244
LLaDA2.0: Scaling Up Diffusion Language Models to 100B Paper • 2512.15745 • Published Dec 10, 2025 • 89
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Paper • 2504.01990 • Published Mar 31, 2025 • 305