NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation Paper • 2606.03159 • Published Jun 2 • 23
Variance Reduction for Expectations with Diffusion Teachers Paper • 2605.21489 • Published May 21 • 1
Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting Paper • 2509.22195 • Published Sep 26, 2025
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty Paper • 2408.14339 • Published Aug 26, 2024
Benchmarking Visual State Tracking in Multimodal Video Understanding Paper • 2606.03920 • Published Jun 2 • 53
Video Models Reason Early: Exploiting Plan Commitment for Maze Solving Paper • 2603.30043 • Published Mar 31 • 14
WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation Paper • 2603.29089 • Published Mar 31 • 9
Test-Time Training with KV Binding Is Secretly Linear Attention Paper • 2602.21204 • Published Feb 24 • 32
Ego4D: Around the World in 3,000 Hours of Egocentric Video Paper • 2110.07058 • Published Oct 13, 2021 • 1
ICONS: Influence Consensus for Vision-Language Data Selection Paper • 2501.00654 • Published Dec 31, 2024
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? Paper • 2410.03859 • Published Oct 4, 2024 • 3
Explain Before You Answer: A Survey on Compositional Visual Reasoning Paper • 2508.17298 • Published Aug 24, 2025 • 4