unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit Image-Text-to-Text • 9B • Updated Oct 31, 2025 • 23.5k • 22
google/siglip2-base-patch16-512 Zero-Shot Image Classification • 0.4B • Updated Feb 21, 2025 • 82.4k • 48
microsoft/Phi-4-multimodal-instruct Automatic Speech Recognition • 6B • Updated Dec 10, 2025 • 538k • 1.61k
microsoft/Phi-3.5-vision-instruct Image-Text-to-Text • 4B • Updated Dec 10, 2025 • 1.24M • 737
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146
huihui-ai/Huihui-Qwen3-VL-4B-Instruct-abliterated Image-Text-to-Text • 4B • Updated Dec 15, 2025 • 22.6k • • 92
meta-llama/Llama-4-Scout-17B-16E-Instruct Image-Text-to-Text • 109B • Updated May 22, 2025 • 449k • • 1.33k