Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published 12 days ago • 93
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Paper • 2604.08539 • Published Apr 9 • 52
bryanzhou008/vit-mae-base-finetuned-eurosat Image Classification • 85.8M • Updated Oct 21, 2024 • 19 • 1
bryanzhou008/swin-tiny-patch4-window7-224-finetuned-eurosat Image Classification • 27.6M • Updated Oct 30, 2024 • 14 • 1
bryanzhou008/vit-base-patch16-224-in21k-finetuned-eurosat Image Classification • 85.8M • Updated Oct 30, 2024 • 7 • 1
bryanzhou008/vit-base-patch16-224-in21k-finetuned-inaturalist Image Classification • 85.8M • Updated Aug 19, 2025 • 73 • 2
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction Paper • 2511.20937 • Published Nov 26, 2025 • 16
Running on Zero Agents Featured 115 SAM3 Video Segmentation 🐠 115 Track and label objects in videos using text prompts or clicks