ExLlamaV3 Major Updates
ExLlamaV3 Major Updates!
ExLlamaV3 added DFlash in v0.0.31, raising Coding throughput from 59.21 t/s to 177.67 t/s; v0.0.32 optimized five models, with Trinity-Nano gaining 72.4% on 6000 Pro², while v0.0.33 adds DFlash model quantization plus bug fixes and efficiency work.
Why it matters: HKR-H/K/R all pass, but the blast radius is mostly LocalLLaMA and ExLlama users. This fits a mid-weight open-source inference update, not a same-day industry-wide story.