Running Reproduction: DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models 🎯 Explore and manage project logs, code, and traces
Running Reproduction: FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations 🎯 Browse project code, traces, and workspace in an interactive logbook
Running Reproduction: Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought 🎯 Explore experiment logs and sync findings with an AI agent
Running Reproduction: OSF: On Pre-training and Scaling of Sleep Foundation Models 🎯 Explore project logs, traces, and workspace in an interactive Logbook
Running Reproduction: SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads? 🎯 Explore code logs, traces, and workspace in a web logbook
Running Reproduction: Online Social Welfare Function-based Resource Allocation 🎯 Explore experiment logs and collaborate with an AI agent
Running Reproduction: Online Social Welfare Function-based Resource Allocation 🎯 Explore and reproduce online resource allocation experiment logs
Running Reproduction: Real-Time Visual Attribution Streaming in Thinking Model 🎯 Explore experiment logs and collaborate with an AI agent
Running Reproduction: Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep Learning 🎯 Collaborate on a research logbook with an AI coding agent
Running Reproduction: Rethinking Genomic Modeling Through Optical Character Recognition 🎯 Explore a collaborative research logbook with AI assistance
Running Reproduction: HECTOR: Hybrid Editable Compositional Object References for Video Generation 🎯 Explore research logbook and collaborate with an AI agent
Running Reproduction: PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models 🎯 Organize visual aesthetic plans and collaborate with an AI agent
Running Reproduction: PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models 🎯 Plan photo aesthetics with AI and log the results
Running Reproduction: No Data? No Problem: Robust Vision-Tabular Learning with Missing Values 🎯 Explore experiment logs and sync findings with an AI agent
Running Reproduction: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? 🎯 Explore a coding task logbook and sync with AI agents
Running Repro - FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations 🎯 Explore and collaborate on a benchmark logbook