A True 2-bit KV Quantization Algorithm for Long-context Reasoning Beyond TurboQuant
超越TurboQuant,面向长上下文推理的真2-bit KV Quantization算法问世
TogetherAI and collaborators released OSCAR, a 2.28 BPE INT2 KV Cache system integrated with SGLang, reporting up to 3× decode speedup at 100k context and up to 7× job-level throughput under a fixed memory budget.
Why it matters: HKR-H/K/R pass, but this is niche inference optimization rather than a broad model launch. The 100k-context and ~3×/~7× claims justify a featured score, not same-day must-write.