baohao/ReOPD_Math_Qwen3-8B_SFT-RL_to_Math_Qwen3-4B-Instruct-2507_SFT 4B • Updated 11 days ago • 30 • 1
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish Paper • 2606.18717 • Published Jun 17 • 6
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 172
Representation Forcing for Bottleneck-Free Unified Multimodal Models Paper • 2605.31604 • Published May 29 • 63
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps Paper • 2605.16928 • Published May 16 • 99