Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Yi Cui's picture

Yi Cui

onekq
24 20 1
Abrahrm's profile picture clippinglab071's profile picture nhlayisekobvuma's profile picture
ยท
  • onekq_ai
  • onekq
  • yicui

AI & ML interests

Benchmark, Code Generation Model

Recent Activity

posted an update about 18 hours ago
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too). MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K. On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX. For high-end inferencing, I place my bet on workstations over Macs.
posted an update 4 days ago
My biggest takeaway from Ben Thompson interview is the legacy of this round of "bubble". For railroad, it was the track, and dot com, fiber. Compared to these two, GPUs depreciate too fast. The answer is power plants.
posted an update 5 days ago
NVidia?!
View all activity

Organizations

MLX Community's profile picture ONEKQ AI's profile picture CSC Generation's profile picture
onekq 's papers 4
arxiv:2505.09027
arxiv:2409.13773
arxiv:2409.05177
arxiv:2408.00019
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs