view article Article Wire It, Run It, Deploy It: AI Workflows in Gradio ysharma, abidlabs • 4 days ago • 24
Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving Paper • 2605.18137 • Published May 27 • 1
view article Article Compressing Time: A Comparative Study of Video VAEs in Diffusers Bekhouche • May 28 • 2
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images Paper • 2412.08802 • Published Dec 11, 2024 • 7
V-JEPA 2 Collection A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of https://ai.meta.com/blog/v-jepa-yann • 8 items • Updated Jun 13, 2025 • 229
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild Paper • 2603.17187 • Published Mar 17 • 141
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Paper • 2502.14786 • Published Feb 20, 2025 • 169
view article Article SigLIP 2: A better multilingual vision language encoder +1 ariG23498, merve, qubvel-hf • Feb 21, 2025 • 226