Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Yi Cui
onekq
24
20
1
Follow
tuiko's profile picture
sleepygargoyle's profile picture
Moksha2024's profile picture
248 followers
ยท
30 following
onekq_ai
onekq
yicui
AI & ML interests
Benchmark, Code Generation Model
Recent Activity
posted
an
update
about 5 hours ago
CoT: ๐ฌ๐ฌ๐ฌ๐ฌ Astra: ๐ง ๐ง ๐ง ๐ฌ
replied
to
their
post
1 day ago
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too). MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K. On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX. For high-end inferencing, I place my bet on workstations over Macs.
replied
to
their
post
3 days ago
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too). MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K. On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX. For high-end inferencing, I place my bet on workstations over Macs.
View all activity
Organizations
Posts
124
view post
Post
29
CoT: ๐ฌ๐ฌ๐ฌ๐ฌ
Astra: ๐ง ๐ง ๐ง ๐ฌ
View all Posts
Articles
8
Article
๐ Matrices in Transformers: Preface
View all Articles
Papers
4
arxiv:
2505.09027
arxiv:
2409.13773
arxiv:
2409.05177
arxiv:
2408.00019
models
0
None public yet
datasets
0
None public yet