Nvidia is targeting the ballooning cost of always-on AI agents with a fast, open 30-billion-parameter model and a routing tool that sends different parts of a task to models based on capability, cost and latency.
The chipmaker on Tuesday introduced Nemotron 3.5 Lightning, an open AI model, and NeMo Switchyard, an open-source routing library aimed at companies building autonomous AI agents across local systems, data centers and cloud environments.
Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model with 3 billion active parameters, designed for high-volume tasks inside larger agent systems. Nvidia said the model delivers up to four times the output speed of similar-sized models. On PinchBench, it achieved 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at comparable accuracy.
The release marks Nvidia’s first open-source AI model since CEO Jensen Huang publicly argued that open-weight AI models should remain widely available in the United States.
Routing requests to the right model
Alongside the model, Nvidia released NeMo Switchyard, an open-source routing library that can send requests, workflow stages or sessions to different models based on factors such as capability, latency and cost.
The company said enterprises can use Switchyard with mixtures of open, proprietary and Nvidia models without rewriting existing applications. Internal benchmarks cited by Nvidia show the software can maintain frontier-level accuracy while reducing task completion costs to roughly one-third of using Anthropic’s Opus 4.8 model alone.
Nvidia said companies including CrowdStrike, Harvey, CodeRabbit, Ramp, LangChain and Cognition have tested or integrated the new tools into their AI workflows.
Why Nvidia is pushing open AI
The announcement comes as competition around open-source AI intensifies. Meta released a new open coding model this week, while Chinese AI developers have rapidly narrowed the performance gap with leading U.S. systems.
Nvidia has increasingly positioned itself as a major advocate for open models. “Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty,” Huang wrote in a recent public statement.
The company’s business incentive is also clear: free AI models still need powerful hardware to run. As Huang told Axios last month, “Free AI should be great for hardware. Free AI should be great for chips.”
The tradeoff for enterprises
The practical shift is less about finding a single best model and more about building an AI system that knows when to use each model.
That could make Switchyard valuable for companies operating complex agents, particularly where routine requests make up a large share of workloads. But Nvidia's performance and cost figures are largely based on its own benchmarks or partner testing, so buyers should validate savings and accuracy against their own workloads before replacing existing models.
Lightning is not positioned as a frontier general-purpose model. Its advantage is efficient, customizable execution, making the distinction between its 30 billion total parameters and 3 billion active parameters especially important.
Nemotron 3.5 Lightning is available through Hugging Face, ModelScope, OpenRouter and Nvidia's platforms, while NeMo Switchyard is available on GitHub.
Read more: Learn how AI agents plan, use tools and coordinate tasks, and how Nvidia’s new releases could make those workflows faster and less expensive.


