Skip to content
Trending storyPast story

NVIDIA Nemotron 3.5 Lightning: a 30B sparse model built for high-frequency agent execution

1 report1 sourceupdated 8 days ago

What happened

Summary

Nemotron 3.5 Lightning 是个 30B 参数的混合专家模型,但每次推理只激活大约 3B 参数,所以跑得快、成本低。NVIDIA 把它设计给 Agent 流程里的执行层用:比如选工具、读文件、检查结果这些需要反复调用模型、但单次任务很明确的步骤。它跟 Nemotron 3 Ultra 不是替代关系,Ultra 负责费脑子的规划,Lig...

Coverage

Follow the reports to see the story from different sides.

Sep 22
  1. AI HOT (Curated Pool)Pick
    NVIDIA Nemotron 3.5 Lightning: a 30B sparse model built for high-frequency agent execution

    NVIDIA positions Nemotron 3.5 Lightning as the execution layer in agent workflows—handling frequent tool calls, file reads, and result checks rather than heavy planning. It's a 30B MoE model that activates only ~3B parameters per token, keeping latency and cost low for high-volume calls. It complements, not replaces, Nemotron 3 Ultra. Weights are open, with tool calling and structured output support. Context goes up to 1M tokens, though OpenRouter's standard tier caps at 262K. Worth a look if your agent makes many model calls per run.