Nvidia released Nemotron 3.5 Lightning on Tuesday, a 30-billion-parameter mixture-of-experts model with 3 billion active parameters, free to download and pitched at the autonomous-agent workloads that enterprises are quietly betting the next capex cycle on. Alongside it: NeMo Switchyard, an open-source routing library that, per Nvidia, “can intelligently direct each request to the most capable and suitable model for the job, across developers’ own mix of open, proprietary and NVIDIA models, without requiring developers to rewrite their applications.”

It’s Nvidia’s first open-source model since Jensen Huang broke roughly three weeks of X silence to argue that open weights were, in fact, good for the company that sells the shovels. A month ago he told Axios: “Free AI should be great for hardware. Free AI should be great for chips.” The release is that thesis rendered as a product.

The competitive backdrop is unsubtle. Moonshot AI’s Kimi K3 narrowed the gap with the frontier American labs earlier this year and set off a Washington panic about distillation and national security. This week Mark Zuckerberg published a lengthy manifesto making his own open-source case, and Meta shipped a coding model called Muse Spark. Nvidia is now boxed in on both flanks: Chinese open weights it can’t ignore, and a Meta that’s aggressively narrative-managing itself into the role of the West’s open-source standard-bearer. Answering with a model of its own is the only move that keeps the demand curve for GPUs pointed the right way.

The engineering claims are aimed squarely at that argument. Nvidia says Lightning delivers up to 4x the output speed of comparable models and completes agentic tasks 30% faster, all while running on a single GPU across Jetson, RTX, and DGX Spark. Kari Briski, Nvidia’s VP of Generative AI, offered the demo statistic the sales org will be repeating for months: CodeRabbit trained a router agent for $85 in around two hours, one epoch, using the standard auto recipe. Another partner, she said, dropped Lightning into an existing post-training stack “with no changes required.” Enterprise customers like CrowdStrike and Harvey are the intended audience, less interested in benchmark supremacy than in the arithmetic of not paying OpenAI or Anthropic per token.

The subtext is that open-weight economics only threaten the labs whose moat was the API. For the company selling the silicon underneath every deployment, free models are demand generation.

Sources