Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries Paper • 2610.06269 • Published 6 days ago
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research Paper • 2610.02202 • Published 10 days ago • 14
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published Aug 13 • 16
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 56
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published Jul 26 • 97
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published Jul 26 • 129
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Paper • 2604.14142 • Published Apr 15 • 30
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Paper • 2603.24440 • Published Mar 25 • 98
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation Paper • 2603.18886 • Published Mar 19 • 6
Grounding and Enhancing Informativeness and Utility in Dataset Distillation Paper • 2601.21296 • Published Jan 29 • 21
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence Paper • 2511.18538 • Published Nov 23, 2025 • 307
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique Paper • 2511.09067 • Published Nov 12, 2025 • 2
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models Paper • 2510.27196 • Published Oct 31, 2025
Grounding Computer Use Agents on Human Demonstrations Paper • 2511.07332 • Published Nov 10, 2025 • 107