Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published 4 days ago • 139
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting Paper • 2607.28261 • Published 8 days ago • 116
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 8 days ago • 302
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published 24 days ago • 20
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Paper • 2607.06987 • Published about 1 month ago • 9
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Paper • 2605.12034 • Published May 13 • 6
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization Paper • 2605.13641 • Published May 13 • 51