So, what are you going to do with these findings? Will you need to get Apple involved? Will you submit pull requests to the torch or mlx folks? Any chance of further improvements with more processing time? Can the regressions be mitigated with some kind of switch between the standard kernels and your optimized ones?
Aaron Reitz
Adreitz
·
AI & ML interests
None yet
Recent Activity
new activity 10 days ago
LAXMAYDAY/FLUX.2-dev-int8-tensorwise:Benchmark vs Q8 GGUF? new activity 23 days ago
Qwen/Qwen3.8-27B:We Cracked Qwen3.8-27B Quant: 27GB INT4 that actually thinks (Heretic Edition)Organizations
None yet