FunAudioLLM/Fun-ASR-Nano-2512-hf Automatic Speech Recognition • 0.8B • Updated 4 days ago • 13.3k • 13
KVAE: Family of Tokenizers for Multimodal Generative Models Paper • 2608.05798 • Published 10 days ago • 28
Kandinsky WM 1.0 Collection Image-to-Video models for Physical AI: autonomous driving, robotics, general physics. • 3 items • Updated 11 days ago • 3
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Paper • 2606.30642 • Published Jun 29 • 1