NVIDIA Vera Rubin NVL72 claims up to 30x more work per watt for AI agents
NVIDIA Vera Rubin NVL72 树立 AI 智能体效率新标准:每瓦特工作量提升至 30 倍
NVIDIA claims the Vera Rubin NVL72 rack-scale system delivers up to 30x more work per watt than H100 when running AI agent inference. The internal test used Llama 3.3 70B on agent workflows. The post doesn't disclose latency figures or full test configs, so treat 30x as a peak number. The system is slated for second-half 2026 and targets inference and agent workloads.
Why it matters: NVIDIA's 30x efficiency claim comes with a concrete model and scenario, not just marketing fluff. But the post doesn't disclose latency or full test configs, so real-world performance will be lower. Useful for infra folks, less resonant for general AI practitioners.