SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 6 days ago • 96
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback Paper • 2608.30241 • Published 7 days ago • 12
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions Paper • 2608.30428 • Published 7 days ago • 14
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 9 days ago • 27
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 7 days ago • 141
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 13 days ago • 148
Meta$^n$: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 13 days ago • 15
AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI Paper • 2511.20686 • Published Nov 20, 2025
Meta^n: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 13 days ago • 15
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 14 days ago • 11
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 14 days ago • 205
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 18 days ago • 65