Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 150
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
awrenn53/gr00t-n17-libero-goal-oss-h16-adamw-gbs640-s42-head9adea3f9-012000-20260706 3B • Updated Jul 6 • 1 • 1
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents Paper • 2605.29534 • Published May 28 • 15
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering? Paper • 2605.28255 • Published May 27 • 1
SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent Paper • 2605.24468 • Published May 23 • 9