Key takeaways

  • Nvidia's new Nemotron 3.5 Lightning is a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks with a…
  • 5 Lightning is a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks with a quarter of the parameters…
  • That puts Lightning on par with OpenAI's gpt-oss-120b (24) and just behind Nvidia's own Nemotron 3 Super (26), which is about four times…

What happened

5 Lightning is a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks with a quarter of the parameters while delivering the fastest inference speeds in its class. 5 lineup. 6 billion active at any given time. According to the independent benchmarking platform Artificial Analysis, the model scores 24 on the Intelligence Index, a nine-point jump from its predecessor (15).

Why it matters

That puts Lightning on par with OpenAI's gpt-oss-120b (24) and just behind Nvidia's own Nemotron 3 Super (26), which is about four times larger. 6 35B A3B (32) and Meta's new Muse Glimmer (35), still hold a clear lead. Nvidia is targeting a different spot on the efficiency frontier with Lightning. 5 Flash-Lite (386 tokens/s). 8 minutes. Proprietary models still dominate the overall efficiency frontier.

6 Luna (max) reaches 52 points in under two minutes. The biggest improvements show up in agentic benchmarks, according to Artificial Analysis. On GDPval-AA v2, Lightning reaches an Elo rating of 824, a 334-point gain over Nemotron 3 Nano. That beats both gpt-oss-120b (800) and the larger Nemotron 3 Super (698). 2 percent. 1 license, positioning it as a high-throughput workhorse for agent-based pipelines.

Artificial Analysis reports that Nvidia worked with partners like CodeRabbit and Harvey on post-training to boost performance in specific domains. Nvidia provides the model in both BF16 and NVFP4 weights. The NVFP4 variant also scores 24 on the Intelligence Index with minimal quality loss compared to the higher-precision version, according to Artificial Analysis. The reasoning model handles text only and supports a context window of one million tokens.

What to watch

Weights are available now, and serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others. Nvidia's push for efficiency over size isn't new. In a widely discussed paper last year, its researchers argued that models under 10 billion parameters can handle most agent workloads as well as 70- to 175-billion-parameter models at one-tenth to one-thirtieth the cost.

6 billion per step, putting it in the same lightweight class. At nearly 670 tokens per second, it also beats gpt-oss-120b and the larger Nemotron 3 Super on agentic benchmarks, making it the clearest product-level proof of that thesis yet.