wenhua cheng
wenhuach
AI & ML interests
Model Compression, CV
Recent Activity
liked a model 1 day ago
Intel/GLM-5.3-Flash-W4A16-AutoRound liked a model 1 day ago
Intel/Qwen3.8-Flash-Next-W4A16-AutoRound liked a dataset 2 days ago
allenai/dolmaOrganizations
AutoRound has recently supported AWQ+SignRound
👍 1
#1 opened 8 days ago
by
wenhuach
Need help
4
#1 opened 19 days ago
by
tooltd
Convert to OpenVINO
1
#2 opened 18 days ago
by
savvadesogle
spills reasoning into code
2
#3 opened 24 days ago
by
MilestoneAI
This quant is the best. Please conver it to GGUF
5
#7 opened 3 months ago
by
alexcardo
torch RuntimeError: Shape mismatch: a.size(1) = 4096, size_k = 8192
2
#1 opened 3 months ago
by
saadsafi
INT8 version for TP=2 / dual Ampere GPUs?
🚀 1
1
#6 opened 4 months ago
by
mancub
Correct metadata, add library name, and link SignRoundV2 paper
#1 opened 4 months ago
by
nielsr
Why delete Intel/Qwen3.6-35B-A3B-int4-AutoRound?
15
#1 opened 4 months ago
by
bgeneto
it can use dflash directly with z-lab/Qwen3.6-35B-A3B-Dflash
1
#2 opened 4 months ago
by
syvvvv
Please update chat template
2
#4 opened 4 months ago
by
alexcardo
does this even run on intel gpus?
7
#2 opened 4 months ago
by
Thomas98519864
AutoRound quant fails to load with mlx-lm
👍 1
1
#1 opened 4 months ago
by
smcleod
How does this compare to the original 8bit qwen quant and the 4 bit auto-round quant?
2
#5 opened 5 months ago
by
sparx3
any plan for an Ampere compatible version?
2
#2 opened 4 months ago
by
electroglyph
Fails to load on Ampere (sm_86) at TP=2: Marlin kernel rejects 32-dim weight slice
2
#3 opened 5 months ago
by
wasifb
MTP 0 accept rate
2
#4 opened 5 months ago
by
AMUN-RA1
Installation Video and Testing - Step by Step
🚀 3
5
#1 opened 5 months ago
by
fahdmirzac
Performance indicators
👍 4
4
#1 opened 5 months ago
by
dehnhaide