SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 3 days ago • 56
Fara-1.5: Scalable Learning Environments for Computer Use Agents Paper • 2606.20785 • Published Jun 18 • 3
Fara1.5 Collection Collection of Fara1.5 CUA models in three sizes - 4B, 9B and 27B. • 3 items • Updated 2 days ago • 6
Perceiver IO: A General Architecture for Structured Inputs & Outputs Paper • 2107.14795 • Published Jul 30, 2021 • 2
view article Article Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World +3 daniel-treble, whojavumusic, alessia-treble, georg-goetz, bezzam • Jun 24 • 9
Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States Paper • 2606.19334 • Published Jun 17 • 8
SWE-FastContext Collection A family of code-search models powering the Explore subagent for coding agents.(It will be made public later) • 3 items • Updated 25 days ago • 18
view article Article MTEB Leaderboard: From a slow demo to feature-rich leaderboard Samoed • Jun 12 • 22
view article Article Unlocking asynchronicity in continuous batching +1 ror, pcuenq, ariG23498 • May 14 • 65
view article Article Continuous batching from first principles +1 ror, ArthurZ, mcpotato • Nov 25, 2025 • 423
view article Article Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP +3 ariG23498, ror, sergiopaniego, pcuenq, sayakpaul • Jun 11 • 56
view article Article Introducing North Mini Code: Cohere’s First Model For Developers CohereLabs • Jun 9 • 83
view article Article olmo-eval: An evaluation workbench for the model development loop allenai • Jun 12 • 17
Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning Paper • 2605.28424 • Published May 27 • 32
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training Paper • 2605.29888 • Published May 28 • 34
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis Paper • 2605.30434 • Published May 28 • 23