Nemotron 3.5 Lightning
Nemotron 3.5 Lightning (NVIDIA, 2026): what the release is, why it matters for operators, specs, benchmarks, and the call.

--- title: “Nemotron 3.5 Lightning” description: “Nemotron 3.5 Lightning (NVIDIA, 2026): what the release is, why it matters for operators, specs, benchmarks, and the call.” type: guide category: ai pubDate: 2026-08-12 image: /images/abs-model-nemotron-3.5-lightning.png imageAlt: “A retro robot representing Nemotron 3.5 Lightning” imagePrompt: “Bold graphic editorial illustration, 1990s comic-book influence, heavy ink outlines, halftone texture, crimson and electric blue on cream. A single retro-futurist robot representing an AI model / a brain-in-a-server, no text, no logos, 16:9.” affiliate: false sources: - name: “NVIDIA Developer Blog — Nemotron 3.5 Lightning delivers fast, accurate, specialized task execution for long-running agents (Aug 11, 2026)” url: “https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/” - name: “NVIDIA Blog — Nemotron Lightning + Switchyard on RTX and DGX (Aug 11, 2026)” url: “https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/” - name: “Hugging Face model card — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16” url: “https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16” - name: “PureAI — Nvidia’s Nemotron 3.5 Lightning Targets the Workhorse Role in Agentic AI, David Ramel (Aug 11, 2026)” url: “https://pureai.com/articles/2026/08/11/nvidias-nemotron-3-5-lightning-targets-the-workhorse-role-in-agentic-ai.aspx” - name: “The Decoder — Nvidia’s open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence, Matthias Bastian (Aug 11, 2026)” url: “https://the-decoder.com/nvidias-open-weight-nemotron-3-5-lightning-prioritizes-speed-over-maximum-intelligence/” - name: “Artificial Analysis — Nemotron 3.5 Lightning model page and launch analysis (Aug 11, 2026)” url: “https://artificialanalysis.ai/models/nemotron-3-5-lightning” - name: “Futurum — Who decides which model runs? NVIDIA would like a say, Nick Patience (Aug 11, 2026)” url: “https://futurumgroup.com/insights/who-decides-which-model-runs-nvidia-would-like-a-say/” - name: “DataCamp — Nemotron 3.5 Lightning features and benchmarks (Aug 11, 2026)” url: “https://www.datacamp.com/blog/nemotron-3.5-lightning” facts: - label: “Vendor” value: “NVIDIA” - label: “Released” value: “2026” - label: “License” value: “Open-weight (NVIDIA)” - label: “Type” value: “AI model release” related: - abs-model-claude-opus-5 - abs-model-gpt-5.6-cyber - abs-model-grok-4.6 - abs-model-kimi-k3 - abs-model-lfm2.5-vl-3b - abs-model-ling-3.0-flash - abs-model-muse-glimmer-30b - abs-model-muse-spark-1.2 - abs-model-qwen-image-3.0 - abs-model-qwen3.7-flash - abs-model-qwen3.8-2.4t-a95b - abs-model-qwen3.8-max-weights-update - abs-model-solar-pro-4 tags: - nemotron-3.5-lightning - ai-models - model-release draft: false --- --- title: “Nemotron 3.5 Lightning” description: “Nemotron 3.5 Lightning (NVIDIA, 2026): what the release is, why it matters for operators, specs, benchmarks, and the call.” type: guide category: ai pubDate: 2026-08-12 image: /images/abs-model-nemotron-3.5-lightning.png imageAlt: “A retro robot representing Nemotron 3.5 Lightning” imagePrompt: “Bold graphic editorial illustration, 1990s comic-book influence, heavy ink outlines, halftone texture, crimson and electric blue on cream. A single retro-futurist robot representing an AI model / a brain-in-a-server, no text, no logos, 16:9.” affiliate: false sources: - name: “NVIDIA Developer Blog — Nemotron 3.5 Lightning delivers fast, accurate, specialized task execution for long-running agents (Aug 11, 2026)” url: “https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/” - name: “NVIDIA Blog — Nemotron Lightning + Switchyard on RTX and DGX (Aug 11, 2026)” url: “https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/” - name: “Hugging Face model card — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16” url: “https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16” - name: “PureAI — Nvidia’s Nemotron 3.5 Lightning Targets the Workhorse Role in Agentic AI, David Ramel (Aug 11, 2026)” url: “https://pureai.com/articles/2026/08/11/nvidias-nemotron-3-5-lightning-targets-the-workhorse-role-in-agentic-ai.aspx” - name: “The Decoder — Nvidia’s open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence, Matthias Bastian (Aug 11, 2026)” url: “https://the-decoder.com/nvidias-open-weight-nemotron-3-5-lightning-prioritizes-speed-over-maximum-intelligence/” - name: “Artificial Analysis — Nemotron 3.5 Lightning model page and launch analysis (Aug 11, 2026)” url: “https://artificialanalysis.ai/models/nemotron-3-5-lightning” - name: “Futurum — Who decides which model runs? NVIDIA would like a say, Nick Patience (Aug 11, 2026)” url: “https://futurumgroup.com/insights/who-decides-which-model-runs-nvidia-would-like-a-say/” - name: “DataCamp — Nemotron 3.5 Lightning features and benchmarks (Aug 11, 2026)” url: “https://www.datacamp.com/blog/nemotron-3.5-lightning” facts: - label: “Vendor” value: “NVIDIA” - label: “Released” value: “2026” - label: “License” value: “Open-weight (NVIDIA)” - label: “Type” value: “AI model release” related: - abs-model-claude-opus-5 - abs-model-gpt-5.6-cyber - abs-model-grok-4.6 - abs-model-kimi-k3 - abs-model-lfm2.5-vl-3b - abs-model-ling-3.0-flash - abs-model-muse-glimmer-30b - abs-model-muse-spark-1.2 - abs-model-qwen-image-3.0 - abs-model-qwen3.7-flash - abs-model-qwen3.8-2.4t-a95b - abs-model-qwen3.8-max-weights-update - abs-model-solar-pro-4 tags: - nemotron-3.5-lightning - ai-models - model-release draft: false --- ## The problem Long-running AI agents spend most of their tokens on the boring work. Pull a file, validate a JSON response, format an output, dispatch a sub-task, retry a tool call. None of that needs a frontier reasoning model — but most stacks route every call through one anyway, because there was never a serious open-weight alternative. NVIDIA built Nemotron 3.5 Lightning to fix that. As the launch team put it directly: “Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning model for every execution step adds cost and latency” — Chris Alexiuk & Chintan Patel, NVIDIA Developer Blog, Aug 11, 2026. The fix is a model that is small where it counts, fast where it matters, and purpose-trained for the harnesses an operator already runs. ## What it is Nemotron 3.5 Lightning is an open 30B / 3B-active hybrid MoE — Mixture of Experts, an architecture where only a small subset of the total parameters runs on any given token — released on August 11, 2026 under the permissive OpenMDW-1.1 license. NVIDIA distilled it from their frontier Nemotron 3 Ultra (~550B / 55B) over roughly six weeks. Total parameters: 30B (31.6B per The Decoder). Active parameters per token: 3B. The “Lightning” name signals the tier it occupies: highest throughput, lowest latency, the call-and-response work that dominates an agent’s token budget. The model sits inside a deliberate three-tier family. Ultra is the frontier reasoning orchestrator (550B / 55B). Super is the middle (120B / 12B, AA Intelligence Index 26). Lightning is “the workhorse layer” — high-volume tool calls, validation, formatting, always-on agent execution. Per PureAI: “Nvidia has released Nemotron 3.5 Lightning, a new open AI model designed less as an all-purpose answer engine than as a fast workhorse inside long-running AI agent systems.” It runs on real operator hardware: a single H100 80GB (validated at 256K context, scaling to 1M on GB200/B200 or 8×H100), a DGX Spark, an RTX 5090, or even Jetson. Day-zero availability covers Hugging Face (BF16 reference, NVFP4 deploy, DSpark-tuned), ModelScope, NVIDIA NIM (NVIDIA Inference Microservice — a hosted OpenAI-compatible API endpoint), OpenRouter, and ten hosted-inference providers including DeepInfra, Fireworks, Together, and Modal. Local tools (LM Studio, llama.cpp, Ollama, Unsloth, Exo, Canonical) all work. ## How it works The model is a hybrid MoE — interleaved Mamba-2 (a state-space sequence architecture with linear-time compute) + MoE + selective Attention layers — succeeding the hybrid Mamba-Transformer design from Nemotron 3 Nano. A “latent MoE” trick doubles the effective expert count at inference. Context length is up to 1M tokens; single H100 deploys are validated at 256K. It is text-only (English, Spanish, French, German, Italian, Japanese — no vision, audio, or video). The reasoning mode is configurable on or off via the chat template (enable_thinking=True/False), with a 16,384-token reasoning budget. The serving stack is where Lightning actually wins: - Multi-token prediction (MTP). The model predicts several tokens at once per forward pass, doubling effective throughput for the validation and formatting calls that dominate agent workloads. - Two speculative-decoding drafts. DSpark (DSpark draft) is tuned for the DGX Spark / low-concurrency case; DFlash draft (DFlash draft) covers broader serving. Speculative decoding means a small draft model proposes tokens the main model then verifies in parallel, which sharply cuts latency. - NVFP4 quantization. A 4-bit NVIDIA floating-point format that holds quality. The model scores 24 on the Artificial Analysis Intelligence Index — an independent composite benchmark that averages across math, reasoning, knowledge, and coding tests — in both BF16 and NVFP4 with negligible quality loss. (AA Intelligence Index is the field’s standard composite for “how smart is this model overall.”) - Harness-optimized training. NVIDIA explicitly tuned the model for popular agent harnesses so tool calls are more accurate with less latency. Per the launch blog: “It is designed for harnesses like OpenClaw and Hermes Agent—all supported by the NVIDIA NemoClaw open source security and management stack for running always-on AI agents” — NVIDIA Developer Blog, Aug 11, 2026. The headline performance numbers: ~670 tokens/second on the NVFP4 prerelease endpoint — the fastest in its size class — and up to 4× faster than similar-sized peers. AA Intelligence Index lands at 24, tying gpt-oss-120b and gaining nine points from predecessor Nemotron 3 Nano (15). On PinchBench — a 10,000-task agent benchmark designed to measure execution reliability rather than raw intelligence — Lightning hits 86% accuracy (NVIDIA) / 85.37 (BF16) and completes the suite 30% faster than Qwen3.6 35B at matched accuracy. The shipping pattern is “frontier reasoning for planning + Lightning for execution, routed by Switchyard.” NVIDIA released NeMo Switchyard alongside the model — an open-source Rust proxy that translates between OpenAI Chat, Anthropic Messages, and OpenAI Responses formats and routes requests across any OpenAI-compatible endpoint (vLLM, NIM, Ollama). A Hermes Agent operator can stand up a two-tier router — Ultra or Super for planning, Lightning for execution — and Switchyard handles the protocol translation invisibly. NVIDIA’s own benchmarks: adding Lightning alongside Opus cut task cost from ~$180 to ~$95 with completion essentially flat. Almost all the benefit came from the first substitution; a two-model pool is the right starting point. ## Where it works Lightning is the right model for the parts of an agent that already work and just need to be cheap: - High-volume tool calls in a long-running agent. Always-on agents doing dozens of tool calls per task route 70–90% of those calls to a 3B-active local model hitting 670 tok/s. For an operator running on
Sources
- NVIDIA Developer Blog — Nemotron 3.5 Lightning delivers fast, accurate, specialized task execution for long-running agents (Aug 11, 2026)
- NVIDIA Blog — Nemotron Lightning + Switchyard on RTX and DGX (Aug 11, 2026)
- Hugging Face model card — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
- PureAI — Nvidia’s Nemotron 3.5 Lightning Targets the Workhorse Role in Agentic AI, David Ramel (Aug 11, 2026)
- The Decoder — Nvidia’s open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence, Matthias Bastian (Aug 11, 2026)
- Artificial Analysis — Nemotron 3.5 Lightning model page and launch analysis (Aug 11, 2026)
- Futurum — Who decides which model runs? NVIDIA would like a say, Nick Patience (Aug 11, 2026)
- DataCamp — Nemotron 3.5 Lightning features and benchmarks (Aug 11, 2026)
Sources
- NVIDIA Developer Blog — Nemotron 3.5 Lightning delivers fast, accurate, specialized task execution for long-running agents (Aug 11, 2026)
- NVIDIA Blog — Nemotron Lightning + Switchyard on RTX and DGX (Aug 11, 2026)
- Hugging Face model card — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
- PureAI — Nvidia's Nemotron 3.5 Lightning Targets the Workhorse Role in Agentic AI, David Ramel (Aug 11, 2026)
- The Decoder — Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence, Matthias Bastian (Aug 11, 2026)
- Artificial Analysis — Nemotron 3.5 Lightning model page and launch analysis (Aug 11, 2026)
- Futurum — Who decides which model runs? NVIDIA would like a say, Nick Patience (Aug 11, 2026)
- DataCamp — Nemotron 3.5 Lightning features and benchmarks (Aug 11, 2026)



Submit a take
Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.