Skip to content
AI HOT (Curated Pool)

Ant Group's Zhou Jun: A trillion-parameter model burns a Tesla's worth of compute every 15 minutes—his team shifts from token count to token density to cut long-context cost from exponential to linear

蚂蚁集团周俊AICon演讲:从Token数量到Token密度,万亿参数模型效率优先

Ant Group VP Zhou Jun laid out the math at AICon: running a trillion-parameter model for 15 minutes costs as much as a Tesla. His team's answer is higher token density, not more tokens. A hybrid linear attention architecture—7 parts Lightning Attention, 1 part MLA—drops 256K long-context cost from exponential to linear, freeing compute for reasoning. The Kpop algorithm separates tool-call tokens from natural-language tokens; combined with chain-of-thought pruning and self-distillation, token output shrinks roughly 4× with no capability loss. A 100B-param model beats larger ones on BFCL and other agent benchmarks, small-model flash throughput hits 2.4×, and five-turn conversation cost falls over 10×.

Why it matters: Ant Group VP Zhou Jun's AICon talk offers a concrete architecture for slashing long-context costs, not just hand-waving. But it's a speech recap, not a product launch or open-source release — no real-world results yet — so the score sits right at the featured threshold.

Read the original ↗Export Markdown