Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Wang Weiyi's picture

Wang Weiyi PRO

kaupane
3 24 193
FlameF0X's profile picture rm's profile picture Fishtiks's profile picture
·
  • Mtrya

AI & ML interests

None yet

Recent Activity

reacted to nwaughachukwuma's post with 👍 1 day ago
Can a text-only model + a vision toolkit (mm-ctx) match a native vision model? We benchmarked 4 setups on 23 multimodal tasks (image, video, audio, PDF): • glm-5.2 (text-only) + mm-ctx: 88.4 • gemini-3.5-flash (vision): 83 • deepseek-v4-pro (text-only) + mm-ctx: 79.4 • qwen3.6-35b-a3b (vision): 44.3 The best text-only setup closes 89% of the gap to the top vision model. Full report: https://huggingface.co/blog/vlm-run/text-only-models-with-mm
liked a model 7 days ago
moonshotai/Kimi-K3
liked a Space 9 days ago
OpenMOSS-Team/MOSS-transcribe-diarize
View all activity

Organizations

None yet
kaupane 's papers 1
arxiv:2601.11354
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs