-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 86 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 237 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 162 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
Collections
Discover the best community collections!
Collections including paper arxiv:2608.05987
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 187 -
Mental World Modeling
Paper • 2607.27201 • Published • 61 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 103 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 59
-
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Paper • 2606.30406 • Published • 26 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20 -
Trust Region Policy Distillation
Paper • 2607.04751 • Published • 33 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 104
-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 111 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 102 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Paper • 2607.29209 • Published • 34 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 103 -
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Paper • 2608.05131 • Published • 14 -
On-Policy Self-Distillation without Any Supervision
Paper • 2608.06296 • Published • 180
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 6 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 153 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 13 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
Paper • 2411.11504 • Published • 24 -
Top-nσ: Not All Logits Are You Need
Paper • 2411.07641 • Published • 24 -
Adaptive Decoding via Latent Preference Optimization
Paper • 2411.09661 • Published • 10 -
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
Paper • 2411.13476 • Published • 15
-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 86 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 237 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 162 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 102 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 187 -
Mental World Modeling
Paper • 2607.27201 • Published • 61 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 103 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 59
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Paper • 2607.29209 • Published • 34 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 103 -
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Paper • 2608.05131 • Published • 14 -
On-Policy Self-Distillation without Any Supervision
Paper • 2608.06296 • Published • 180
-
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Paper • 2606.30406 • Published • 26 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 20 -
Trust Region Policy Distillation
Paper • 2607.04751 • Published • 33 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 104
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 6 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 153 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 13 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 111 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
-
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
Paper • 2411.11504 • Published • 24 -
Top-nσ: Not All Logits Are You Need
Paper • 2411.07641 • Published • 24 -
Adaptive Decoding via Latent Preference Optimization
Paper • 2411.09661 • Published • 10 -
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
Paper • 2411.13476 • Published • 15