Nvidia released Nemotron 3.5 Lightning on Tuesday. It is a 30 billion parameter mixture-of-experts model with three billion active at any moment. The architecture is a hybrid of Mamba-2, MoE and attention layers, with a one million token context window. The weights are on Hugging Face and ModelScope.
The licence is the part that matters. It ships under OpenMDW-1.1, which permits commercial use, and Nvidia published the training data and the recipes alongside the weights. It is free for companies to download, use and modify without asking permission or paying Nvidia, CNBC notes.
The speed claim has two numbers
Nvidia leads on output speed of up to four times that of similar-sized models. Read further and the agentic figure is more modest. On PinchBench it hit 86% accuracy while finishing 10,000 tasks 30% faster than Alibaba’s Qwen 3.6 35B at similar accuracy.
Both numbers are in the same release. Four times faster describes token generation in the lab. Thirty percent is what happens when the model actually does a job. That gap is worth carrying whenever a vendor quotes throughput.
The benchmark scores are respectable rather than startling. It posts 81.94 on MMLU Pro, 75.44 on GPQA Diamond and 51.56 on SWE-bench Verified. Pre-training ran to more than 20 trillion tokens.
The router is the real product
Nvidia also released NeMo Switchyard, an open-source library. It sends each step of an agent workflow to whichever model handles it best. Plans route up to a frontier model. Execution routes down to Lightning.
LangChain ran it across 145 tasks, in figures compiled by MarkTechPost. Routing between Lightning and Claude Opus 4.8 cut costs by 74% against using the frontier model alone. Only 7% of calls went to the expensive model, at a cost of roughly six accuracy points.
That is not a pitch to be the best model. It is a pitch to handle 93% of the calls. Nvidia is arguing that most agent work does not need a frontier model. That is a commoditisation argument aimed squarely at the labs.
Why a chip company gives models away
The motive is not mysterious and Nvidia has never hidden it. Cheaper, freely available models mean more inference, and inference runs on GPUs. Commoditising the software expands the market for the hardware underneath it.
We reported the same logic when Nvidia tied its safety work to chip demand. It also puts models behind a signature. Nvidia signed the open-weights letter urging Washington not to crack down. Jensen Huang made his first ever post on X to promote it. OpenAI, Anthropic and Google did not sign.
The trillion-parameter model, and the catch
The Information reports that Nvidia is building Nemotron 4 at more than a trillion parameters. That is up from the 550 billion of Nemotron 3 Ultra. It could be ready as early as late autumn. The company wants it to compete with the best open models in the world.
Here is the line everyone will skip. Even at a trillion parameters, Nemotron 4 would still be smaller than the leading Chinese open models. Moonshot’s Kimi K3 already holds the record as the largest open model ever released.
So the ambition is catch-up, not leadership. Nvidia is building towards a size China passed first, and saying so in its own briefing.
China set this pace and still holds it
The comparison Nvidia chose for its own benchmark was a Chinese model. Lightning goes up against Qwen 3.6 rather than any American open model. At that size, Qwen is the thing to beat.
DeepSeek has spent the year cutting prices on open models until running anything less capable stopped making sense. Alibaba and Moonshot have traded the top of the open charts between them. American open weights have been a policy argument for most of that time rather than a shipping product.
Nemotron 3.5 Lightning changes that in a small way. It is a real model, genuinely open, from the company that sells the shovels.
What to watch instead of the benchmarks
Nvidia has shipped Nemotron models before, including Nemotron Nano Omni for edge agents. None has moved the open-weight centre of gravity away from Chinese labs.
The test is Nemotron 4 and it has a rough date on it. Say it lands in late autumn and tops the open charts. Nvidia will have bought a seat at a table it currently only funds. Say it lands second to Kimi again. The honest description of the strategy is then demand generation for GPUs.
Both readings can be true at once. That is rather the point of giving software away when you sell the hardware.
Get the TNW newsletter
Get the most important tech news in your inbox each week.