nm-testing/DeepSeek-Coder-V2-Lite-Instruct-quantized.w8a8
16B • Updated • 29
nm-testing/l4-scout-int4-debug
109B • Updated • 9
nm-testing/pixtral-12b-FP8-dynamic
Image-Text-to-Text
• 13B • Updated • 523
• 1
nm-testing/TinyLlama-1.1B-Chat-v1.0-W4A16-G128-Asym-Updated-ActOrder
1B • Updated • 14.8k
nm-testing/TinyLlama-1.1B-Chat-v1.0-awq-group128-asym256
1B • Updated • 3
nm-testing/TinyLlama-1.1B-Chat-v1.0-W4A16-G128-Asym-Updated-Channel
1B • Updated • 5
nm-testing/TinyLlama-1.1B-Chat-v1.0-W4A16-G128-Asym-Updated
1B • Updated • 5
nm-testing/Llama-2-7b-hf-gsm8k-quant_w4a16_sym-compressed
7B • Updated • 4
nm-testing/Llama-2-7b-hf-gsm8k-gptq_w4a16_sym-compressed
7B • Updated • 143
nm-testing/Llama-2-7b-hf-gsm8k-awq_w4a16_sym-compressed
7B • Updated • 4
nm-testing/Llama-2-7b-hf-gsm8k-awq_gptq_sym-compressed
7B • Updated • 5
nm-testing/Mixtral-8x7B-Instruct-v0.1-FP8-Dynamic
47B • Updated • 9
nm-testing/Llama-3.1-8B-Instruct-W4A16-G128-shared-pipeline
8B • Updated • 2
nm-testing/Qwen2-VL-2B-Instruct-FP8-dynamic-cli
2B • Updated • 11
nm-testing/Qwen2-VL-2B-Instruct-FP8_DYNAMIC
Image-Text-to-Text
• 2B • Updated • 8
nm-testing/whisper-large-v3-quantized.w4a16
2B • Updated • 6
nm-testing/whisper-large-v3-quantized.w8a8_sq
2B • Updated • 9
nm-testing/whisper-large-v3-quantized.w8a8
2B • Updated • 11
nm-testing/llama2.c-stories110M-gsm8k-fp8_dynamic-compressed
0.1B • Updated • 775
nm-testing/llama2.c-stories110M-gsm8k-recipe_w4a16_actorder_weight-compressed
0.1B • Updated • 814
nm-testing/llama2.c-stories15M
Text Generation
• 24.4M • Updated • 8.86k
nm-testing/Meta-Llama-3-8B-Instruct-FP8-channel-output-activation-kv_cache-qkv_proj
8B • Updated • 3
nm-testing/Meta-Llama-3-8B-Instruct-FP8-channel-output-activation-q_proj
8B • Updated • 7
nm-testing/Meta-Llama-3-8B-Instruct-FP8-channel-output-activation
8B • Updated • 2
nm-testing/Llama-3.2-1B-W4A16-Transforms
5B • Updated • 4
nm-testing/Ministral-8B-Instruct-2410-FP8-dynamic
8B • Updated • 10
nm-testing/TinyLlama-1.1B-Chat-v1.0-W4A16-G128-asym
1B • Updated • 5
nm-testing/Phi-4-mini-instruct-quantized.w4a16.asymmetric
5B • Updated • 7
nm-testing/Qwen1.5-MoE-A2.7B-Chat-quantized.w4a16
14B • Updated • 71.8k
• 1
nm-testing/Moonlight-16B-A3B.w4a16
16B • Updated • 2.04k