Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
π
In a Training Loop
1300.9
TFLOPS
Andy Chen
andynoodles
14
33
Follow
dipankarsarkar's profile picture
1 follower
Β·
24 following
yi-hsiang-chen-tw
AI & ML interests
Feel free to contact me through Linkedin
Recent Activity
replied
to
onekq
's
post
about 9 hours ago
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too). MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K. On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX. For high-end inferencing, I place my bet on workstations over Macs.
liked
a model
about 14 hours ago
zai-org/GLM-5.3
liked
a model
about 14 hours ago
zai-org/GLM-5.3-Flash
View all activity
Organizations
andynoodles
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
21 days ago
Benchmarked on Single GB10 (ASUS GX10)
π
1
1
#5 opened 21 days ago by
andynoodles
New activity in
openai/privacy-filter
about 2 months ago
vllm or something
5
#25 opened 3 months ago by
prudant
New activity in
andynoodles/Taiwan-UrbanPlan
2 months ago
[bot] Conversion to Parquet
#1 opened 2 months ago by
parquet-converter
New activity in
andynoodles/Taiwan-Financial
3 months ago
[bot] Conversion to Parquet
#1 opened 3 months ago by
parquet-converter
New activity in
PaddlePaddle/PaddleOCR-VL-1.6
3 months ago
Repetition penalty recommands
#4 opened 3 months ago by
andynoodles
New activity in
andynoodles/omnidoc-ocr-correction-bench
4 months ago
[bot] Conversion to Parquet
#1 opened 5 months ago by
parquet-converter
New activity in
RedHatAI/DeepSeek-V4-Flash-NVFP4-FP8
4 months ago
2x Nvidia 6000 Pros
3
#2 opened 4 months ago by
mtcl
New activity in
deepseek-ai/DeepSeek-V4-Flash
4 months ago
Unable to run on 2x RTX Pro 6000 (DEEP_GEMM problem)
β
10
17
#15 opened 4 months ago by
stev236
New activity in
RedHatAI/Qwen3.6-35B-A3B-NVFP4
5 months ago
vllm-openai:cu130-nightly Error
β
4
3
#1 opened 5 months ago by
andynoodles
New activity in
Qwen/Qwen3.5-35B-A3B-GPTQ-Int4
5 months ago
Why 35B-INT4 smaller than 27B-INT4
3
#10 opened 6 months ago by
andynoodles
New activity in
PaddlePaddle/PP-DocLayoutV3
6 months ago
PP-DocLayoutV3 Inference Benchmark: SafeTensor vs ONNX vs PaddlePaddle
π
1
#8 opened 6 months ago by
andynoodles
New activity in
Qwen/Qwen3.5-397B-A17B-GPTQ-Int4
6 months ago
Poor performance in vLLM
4
#3 opened 6 months ago by
sinebubble
New activity in
zai-org/GLM-OCR
6 months ago
Requesting Example for Structured Information Extraction via cURL
π
3
2
#10 opened 7 months ago by
andynoodles
New activity in
zai-org/GLM-OCR
7 months ago
When will there be better support for vLLM?
6
#6 opened 7 months ago by
Xiakj
Requesting Example for Structured Information Extraction via cURL
π
3
2
#10 opened 7 months ago by
andynoodles