The Sequence Chat - Issue 912: NVIDIA’s Chris Alexiuk Talks About Nemotron, GPUs and Agentic AI

NVIDIA’s new Nemotron 3.5 Lightning, other Nemotron models, architectures and more.

The Sequence Chat - Issue 912: NVIDIA’s Chris Alexiuk Talks About Nemotron, GPUs and Agentic AI

TL;DR

  • NVIDIA's Nemotron model family has evolved through various stages, including synthetic data generation and hybrid models, to contribute open-weight technology and understand infrastructure needs.
  • Nemotron 3 is released in three tiers (Nano, Super, Ultra) designed to align with specific hardware configurations (single GPU, node, NVL72 rack) for optimal performance and cost-effectiveness.
  • NVIDIA's involvement with open models like Nemotron is strategic, aiming to enable a broad range of developers and researchers rather than competing directly in the model market.
  • Open data is considered crucial for true model auditability, providing transparency into model behavior beyond behavioral analysis.
  • Architectural innovations in Nemotron 3 include interleaving Mamba-2 layers with sparse MoE and a reduced number of attention layers, balancing context recall with efficiency.
  • LatentMoE technology enhances latency and throughput for inference by compressing tokens into a latent space, allowing for more experts at a lower cost.
  • Nemotron is expanding into multimodality with models for vision, speech, and retrieval, adapting encoders to the Nemotron 3 backbone.
  • Nemotron 3.5 Lightning is a 30B MoE model designed specifically for AI agents, handling tasks like tool calls and subagent delegation efficiently and accurately.
  • Alexiuk predicts that 'model routing' will evolve into more general 'orchestration' layers for AI agents.