HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 1 day ago • 43
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 9 days ago • 73
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 13 days ago • 43
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe Paper • 2607.03451 • Published 27 days ago • 34
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Paper • 2606.20515 • Published Jun 18 • 41
Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion Paper • 2606.15236 • Published Jun 16 • 22
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 76
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects Paper • 2605.21572 • Published May 20 • 55
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation Paper • 2605.19833 • Published May 19 • 137
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Paper • 2604.28185 • Published Apr 30 • 92