The model uses a mixture-of-experts architecture with 30 billion total parameters, while activating only 3 billion for each task, reducing the computing power needed to run it, Anadolu agency reported.
***
Nemotron 3.5 Lightning can be deployed on a single supported graphics processing unit GPU system, including Nvidia’s DGX Spark or H100, and has a context window of up to 1 million tokens, according to the company’s official model documentation.
The model is intended for tasks such as code review, tool use, security-alert monitoring and answering customer billing questions within larger AI-agent systems.
Nvidia said Nemotron 3.5 Lightning delivers up to four times faster output and completes agentic tasks 30% faster than other open models in its class, based on the company’s benchmark tests.
The chipmaker released the model’s weights under its OpenMDW 1.1 license, allowing businesses and developers to customize it using their own data and workflows.
Companies including cybersecurity firm CrowdStrike, legal AI provider Harvey and software-review platform CodeRabbit have tested and customized the model, said Nvidia.
Alongside the model, Nvidia introduced NeMo Switchyard, an open-source routing library that automatically directs each AI task to the most capable and cost-efficient model available.
The company said internal tests showed Switchyard preserved frontier-level accuracy while reducing task-completion costs to nearly one-third of those incurred by using Anthropic’s Opus 4.8 alone.
10
Aug


