FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models Paper • 2606.27866 • Published Jun 26 • 1
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices Paper • 2607.10183 • Published Jul 14 • 2
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization Paper • 2606.00079 • Published May 22 • 1
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs Paper • 2507.07145 • Published Jul 9, 2025 • 1
prithivMLmods/CapQwen3.6-27B-BLIP3o-Long-Caption-Distilled Image-Text-to-Text • 27B • Updated Jun 1 • 35 • 8
XORTRON - Criminal Computing Collection Release Quality Xortron Apps & Models • 6 items • Updated Apr 6 • 25