NVIDIA has released Nemotron 3.5 Lightning, its first Nemotron 3.5-family model — an open 30-billion-parameter mixture-of-experts model with only about 3 billion active parameters, built for always-on AI agents.
'Lightning is built for that execution layer of agents that runs around the clock,' NVIDIA explains in a blog post. In LLM-powered agents — systems that independently operate tools, write code or query databases — execution speed usually dominates: how fast can the model deliver the next tokens? That is exactly where Lightning aims to shine: up to 4x faster than comparable open models on agentic workloads, according to NVIDIA.
Technically, Lightning is a 30B MoE model with 3B active parameters and a 1-million-token context window — the largest in its size class. It ships as a full-precision (BF16) release and a quantized NVFP4 variant on Hugging Face, plus NVIDIA NIM, DeepInfra and Ollama. The global release in the NGC catalog was dated August 11, 2026.
The approach follows a broader industry trend: instead of ever-larger dense models, developers are turning to sparse MoE architectures that deliver near-frontier quality while running on a single laptop GPU. Nemotron 3.5 Lightning is part of NVIDIA's new Nemotron family, which spans from Nano-30B to an announced Ultra model with 550 billion parameters.
Observers see it as an attempt to help shape the open-model landscape as Meta (Muse Glimmer), Alibaba (Qwen3.8-27B) and others release similarly compact agent models.




