Nvidia Nemotron 3.5 lightning and Nemo Switchyard

Aug 12, 2026 02:35 AM - 2 hours ago 1

As AI shifts from chatbots to autonomous agents, unfastened models are serving marketplace demands for afloat power complete wherever AI runs and really it’s deployed and evolves.

Today, NVIDIA is expanding its Nemotron 3 exemplary family pinch Nemotron 3.5 Lightning, the highest-efficiency exemplary successful its people for long-running agentic AI workloads. This merchandise follows Nemotron 3 Nano and reflects NVIDIA’s committedness to continually improving unfastened models for greater accuracy and speed. 

Built for specialized tasks wrong larger multi-agent systems, Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, helps create smarter and much businesslike agentic applications.

Also, NVIDIA is releasing NeMo Switchyard, an unfastened root room for smart routing wrong celebrated supplier tools. Enterprises tin usage it to build a router based connected their circumstantial needs. When deployed, NeMo Switchyard tin intelligently nonstop each petition to the astir tin and suitable exemplary for the job, crossed developers’ ain operation of open, proprietary and NVIDIA models, without requiring developers to rewrite their applications.

Together, Nemotron 3.5 Lightning and NeMo Switchyard present greater power complete really AI is deployed, wherever it runs and really efficiently it operates — crossed PCs, workstations, information centers and the cloud.

Nemotron 3.5 Lightning delivers frontier-level intelligence successful a small, customizable unfastened exemplary built for high-volume agentic workflows.

Always-On Agents Need a System of Models 

Modern agentic systems — always-on agents — progressively run arsenic systems of models, aliases exemplary ensembles, pinch different models specialized for different tasks. 

NVIDIA Nemotron unfastened models are designed for this architecture. A frontier reasoning exemplary specified arsenic Nemotron 3 Ultra aliases GPT-5.6 whitethorn scheme and orchestrate a workflow, while smaller specialized models for illustration Nemotron 3.5 Lightning tin execute targeted tasks specified arsenic codification review, instrumentality use, information alert monitoring and answering billing questions.

Powering High-Volume Specialized Tasks With Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is simply a afloat customizable unfastened exemplary built for high-volume tasks powering always-on agents. It was developed pinch contributions from the Nemotron Coalition, whose members provided information methodologies, conclusion package and datasets to thief beforehand the model.

The exemplary delivers up to 4x faster output speed, starring to 30% faster agentic task completion compared pinch different models successful its class. And because it’s unfastened and customizable, Nemotron 3.5 Lightning tin beryllium easy post-trained pinch NVIDIA NeMo connected an organization’s ain domain data, devices and workflows to amended accuracy for specialized tasks.

PinchBench benchmarks show that Nemotron 3.5 Lightning delivers faster agentic task completion pinch frontier-level accuracy compared to different models successful its class.

AI leaders crossed industries are customizing Nemotron 3.5 Lightning for their workloads, including CrowdStrike for cybersecurity, Harvey pinch Trajectory for ineligible services and CodeRabbit with Baseten for codification review, helping amended accuracy for domain-specific agentic tasks. Additionally, Lila Sciences is helping to amended reasoning capabilities for agentic tasks crossed beingness and life sciences, and Fastino Labs customized the exemplary and is seeing starring accuracies for package development, finance and healthcare workloads. 

Enterprises person customized Nemotron 3.5 Lightning to execute starring accuracy for their specialized task successful their agentic workflows.

Nemotron 3.5 Lightning besides gives organizations power complete privateness and deployment. It tin tally connected section AI systems — including NVIDIA RTX PCs, NVIDIA DGX Spark, NVIDIA DGX Station and NVIDIA Jetson — to thief users maximize existing infrastructure investments, aliases standard crossed separator AI devices, NVIDIA RTX PRO workstations, information centers and unreality environments for endeavor usage cases. And Nemotron 3.5 Lightning tin tally locally aliases connected premises for high-volume, specialized tasks that require accelerated responses.

Also, arsenic pinch each Nemotron launch, NVIDIA publishes arsenic overmuch of the training information and techniques arsenic licensing permits, which allows for traceability, auditing and training of different models. Alongside Lightning, NVIDIA is releasing Nemotron-RL-Agentic-Terminal-Pivot, an agentic reinforcement learning dataset utilized to post-train it for coding supplier capabilities.

More Efficient AI Apps With Model Routing 

Some models are amended for coding, immoderate for reasoning, immoderate for lightweight tasks and immoderate are optimized to tally locally for greater privateness and efficiency. If customers trust connected 1 default model, they mightiness either overspend aliases suffer quality; if they negociate routing manually, it becomes integration activity that tin slow down a deployment.

NVIDIA NeMo Switchyard is an unfastened root exemplary routing room for AI agents. The exertion routes prompts to the astir tin and businesslike exemplary for each measurement of an supplier workflow automatically, based connected circumstantial needs. Agent exertion developers tin tune aliases modify the router pinch different routing algorithms to lucifer their priorities, specified arsenic quality, latency and costs requirements. In a strategy of models, enterprises tin create powerful AI agents pinch improved tokenomics. 

Internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion costs to astir one-third of Opus 4.8 alone.

NVIDIA soul benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion costs to astir one-third of Opus 4.8 alone.

NVIDIA is moving pinch partners crossed the AI ecosystem to bring intelligent exemplary routing into the devices and platforms developers already use. 

  • Boomi: Evaluated Switchyard crossed 5 routing capabilities, achieving 100% domain-routing accuracy, sending 59% of postulation to a 5x faster fine-tuned exemplary and reducing later-turn latency by 21%.
  • Cadence: Improved ratio by 9.9% by utilizing the ChipStack AI Super Agent for a general verification usage case.
  • Classmethod: Is moving opencode and Fireworks workloads utilizing NeMo Switchyard internally, pinch first testing showing a 27% costs simplification while maintaining quality.
  • Cognition: ​Integrated the NVIDIA NeMo Switchyard staged router into Devin Desktop for NVIDIA soul use, achieving near-frontier capacity connected FrontierCode Main while reducing mean costs by 28% comparative to routing each requests to a azygous underlying frontier model.
  • Kong: Delivers routing pinch NeMo Switchyard natively done Kong AI Gateway.
  • LangChain: With NeMo Switchyard, achieved 74% little costs successful 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model, astatine a 6% accuracy tradeoff.
  • LiteLLM: Is adding NeMo Switchyard arsenic a plug-in into its proxy furniture truthful developers tin entree these benefits without changing their existing stack.
  • Nous Research: Integrated NeMo Switchyard into Hermes to supply developers pinch an easy-to-configure routing strategy to amended supplier efficiency.
  • Ramp: Used NeMo Switchyard to lucifer a frontier model’s capacity while cutting costs by 58% and runtime by 33% successful Ramp SWE-Bench.
  • Siemens: Is benchmarking to amended ratio successful its Fuse EDA AI Agent.

Nemotron 3.5 Lightning is disposable connected Hugging Face, ModelScope, OpenRouter and build.nvidia.com arsenic an NVIDIA NIM microservice arsenic good arsenic done a wide ecosystem of NVIDIA Cloud Partners, post-training platforms, conclusion platforms and unreality work providers. NeMo Switchyard is disposable connected GitHub and coming to partner platforms soon. 

More