AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Paper • 2608.00155 • Published 7 days ago • 12
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published 4 days ago • 23
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published 4 days ago • 139
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 4 days ago • 57
ExplainBench: Evaluating Code Explanations from Agents Paper • 2607.26451 • Published 9 days ago • 12
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents Paper • 2608.03509 • Published 3 days ago • 20
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 3 days ago • 28
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published 4 days ago • 37
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation Paper • 2608.02287 • Published 4 days ago • 29
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published 7 days ago • 20
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published 8 days ago • 36
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 12 days ago • 99
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Paper • 2607.26784 • Published 9 days ago • 28
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space Paper • 2607.25675 • Published 10 days ago • 65
Harness-G: A Graph-Structured Harness for Search Agents Paper • 2607.27652 • Published 8 days ago • 5
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published 8 days ago • 58