SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 4 days ago • 92
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback Paper • 2608.30241 • Published 5 days ago • 11
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions Paper • 2608.30428 • Published 5 days ago • 14
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 7 days ago • 27
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 5 days ago • 139
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 11 days ago • 148
Meta$^n$: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 11 days ago • 15
AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI Paper • 2511.20686 • Published Nov 20, 2025
Meta^n: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 11 days ago • 15
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 12 days ago • 11
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 12 days ago • 205
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 16 days ago • 65