← Latest papers
🤖 machine learning

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Nemotron 3 Nano Omni is a new open-source multimodal model that natively supports audio, text, image, and video inputs, offering improved accuracy, agentic capabilities, and efficient inference through a 30B-A3B backbone and token-reduction techniques, with released checkpoints and code to support further research.

Original authors: NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko, Tuomas Rintamaki, Matthieu Le, Tyler Poon, Danial Mohseni Taheri, Ilia Karmanov, Guilin Liu, Jarno Seppanen, Arushi Goel, Mike Ranzinger, Greg
Published 2026-04-29
📖 5 min read🧠 Deep dive

Original authors: NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko, Tuomas Rintamaki, Matthieu Le, Tyler Poon, Danial Mohseni Taheri, Ilia Karmanov, Guilin Liu, Jarno Seppanen, Arushi Goel, Mike Ranzinger, Greg Heinrich, Guo Chen, Lukas Voegtle, Philipp Fischer, Timo Roman, Karan Sapra, Collin McCarthy, Shaokun Zhang, Fuxiao Liu, Hanrong Ye, Yi Dong, Mingjie Liu, Yifan Peng, Piotr Zelasko, Zhehuai Chen, Nithin Rao Koluguri, Nune Tadevosyan, Lilit Grigoryan, Ehsan Hosseini Asl, Pritam Biswas, Leili Tavabi, Yuanhang Su, Zhiding Yu, Peter Jin, Alexandre Milesi, Netanel Haber, Yao Xu, Sarah Amiraslani, Nabin Mulepati, Eric Tramel, Jaehun Jung, Ximing Lu, Brandon Cui, Jin Xu, Zhiqi Li, Shihao Wang, Yuanguo Kuang, Shaokun Zhang, Huck Yang, Boyi Li, Hongxu Yin, Song Han, Pavlo Molchanov, Adi Renduchintala, Charles Wang, David Mosallanezhad, Soumye Singhal, Luis Vega, Katherine Cheung, Sreyan Ghosh, Yian Zhang, Alexander Bukharin, Venkat Srinivasan, Johnny Greco, Andre Manoel, Maarten Van Segbroeck, Suseella Panguliri, Rohit Watve, Divyanshu Kakwani, Shubham Pachori, Jeffrey Glick, Radha Sri-Tharan, Aileen Zaman, Khanh Nguyen, Shi Chen, Jiaheng Fang, Qing Miao, Wenfei Zhou, Yu Wang, Zaid Pervaiz Bhat, Varun Praveen, Arihant Jain, Ramanathan Arunachalam, Tomasz Kornuta, Ashton Sharabiani, Amy Shen, Wei Huang, Yi-Fu Wu, Ali Roshan Ghias, Huiying Li, Brian Yu, Nima Tajbakhsh, Chen Cui, Wenwen Gao, Li Ding, Terry Kong, Manoj Kilaru, Anahita Bhiwandiwalla, Marek Wawrzos, Daniel Korzekwa, Pablo Ribalta, Grzegorz Chlebus, Besmira Nushi, Ewa Dobrowolska, Maciej Jakub Mikulski, Kunal Dhawan, Steve Huang, Jagadeesh Balam, Yongqiang Wang, Nikolay Karpov, Valentin Mendelev, George Zelenfroynd, Meline Mkrtchyan, Qing Miao, Omri Almog, Bhavesh Pawar, Rameshwar Shivbhakta, Sudeep Sabnis, Ashrton Sharabiani, Negar Habibi, Geethapriya Venkataramani, Pamela Peng, Prerit Rodney, Serge Panev, Richard Mazzarese, Nicky Liu, Michael Fukuyama, Andrii Skliar, Roger Waleffe, Duncan Riach, Yunheng Zou, Jian Hu, Hao Zhang, Binfeng Xu, Yuhao Yang, Zuhair Ahmed, Alexandre Milesi, Carlo del Mundo, Chad Voegele, Zhiyu Cheng, Nave Assaf, Andrii Skliar, Daniel Afrimi, Natan Bagrov, Ran Zilberstein, Ofri Masad, Eugene Khvedchenia, Natan Bagrov, Borys Tymchenko, Tomer Asida, Daniel Afrimi, Parth Mannan, Victor Cui, Michael Evans, Katherine Luna, Jie Lou, Pinky Xu, Guyue Huang, Negar Habibi, Michael Boone, Pradeep Thalasta, Adeola Adesoba, Dina Yared, Christopher Parisien, Leon Derczynski, Shaona Ghosh, Wes Feely, Micah Schaffer, Radha Sri-Tharan, Jeffrey Glick, Barnaby Simkin, George Zelenfroynd, Tomasz Grzegorzek, Rishabh Garg, Aastha Jhunjhunwala, Sergei Kolchenko, Farzan Memarian, Haran Kumar, Shiv Kumar, Isabel Hulseman, Anjali Shah, Kari Briski, Padmavathy Subramanian, Joey Conway, Udi Karpas, Jane Polak Scowcroft, Annie Surla, Shilpa Ammireddy, Ellie Evans, Jesse Oliver, Tom Balough, Chia-Chih Chen, Sandip Bhaskar, Alejandra Rico, Bardiya Sadeghi, Seph Mard, Katherine Cheung, Meredith Price, Laya Sleiman, Saori Kaji, Wesley Helmholz, Wendy Quan, Michael Lightstone, Jonathan Cohen, Jian Zhang, Oleksii Kuchaiev, Boris Ginsburg, Jan Kautz, Eileen Long, Mohammad Shoeybi, Mostofa Patwary, Oluwatobi Olabiyi, Andrew Tao, Bryan Catanzaro, Udi Karpas

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a super-smart assistant who has spent their whole life reading books and looking at pictures. They are brilliant at text and images, but if you tried to talk to them or show them a video with sound, they would be confused.

Nemotron 3 Nano Omni is NVIDIA's new version of this assistant. It's like giving that assistant a pair of ears and a video camera, while also teaching them how to process information much faster and more efficiently.

Here is a simple breakdown of what makes this new model special, using everyday analogies:

1. The "All-in-One" Upgrade

Before this, the assistant (Nemotron Nano V2) could only read text and look at static pictures. The new Nemotron 3 Nano Omni is "omni-modal," which is a fancy way of saying it can handle everything at once: text, images, video, and audio.

  • The Analogy: Think of the old model as a librarian who can only read books. The new model is a librarian who can read books, watch movies, listen to podcasts, and then tell you how all those things connect.

2. The "Brain" and the "Senses"

The model is built on a very efficient brain called Nemotron 3 Nano 30B-A3B. This brain is a "Mixture of Experts" (MoE), which is like a team of specialists where only the right experts wake up to solve a specific problem, saving energy.

  • The Senses: To handle the new inputs, they attached two new "sensors":
    • Vision: A high-tech camera system (C-RADIOv4-H) that looks at images and videos.
    • Hearing: A sophisticated ear (Parakeet-TDT) that understands speech, music, and environmental sounds.

3. Smarter Ways to Look and Listen

The paper highlights three clever tricks the model uses to not get overwhelmed by too much data:

  • Dynamic Resolution (The Flexible Lens): Instead of squishing every photo into a tiny, fixed square (which loses detail), the model adjusts its view like a camera zooming in or out to keep the picture's natural shape.
  • Video Compression (The Time-Lapse): When watching a video, the model doesn't look at every single frame if they are identical. It uses a "3D Convolution" trick to merge pairs of frames, effectively cutting the number of "frames" it needs to process in half.
  • Audio Chunks (The Sound Bite): Instead of listening to a 2-hour podcast as one giant block of noise, it breaks the audio into 30-second clips, making it easier to understand the flow of conversation.

4. The "Training Camp" (How it Learned)

You can't just plug in a camera and expect the model to understand movies immediately. The team used a staged training recipe, like a student progressing through school:

  • Stage 0-1: First, they taught the model to understand pictures and text together.
  • Stage 2-3: Next, they taught it to listen to audio and speak.
  • Stage 4-6: Finally, they put it all together. They fed it massive amounts of data (over 466 billion "tokens" or words) that included long documents, hours of video, and complex conversations.
  • The "Reinforcement Learning" Phase: After the initial training, they played a game of "Try, Fail, Succeed." They gave the model problems, checked if it got the answer right, and rewarded it for thinking logically. This is similar to a coach reviewing game tape with an athlete to improve their strategy.

5. Speed and Efficiency

One of the biggest claims in the paper is that this model is incredibly fast.

  • The Analogy: If other models of similar size are like a sports car that gets stuck in traffic, Nemotron 3 Nano Omni is a high-speed train. It uses special techniques to skip unnecessary steps, allowing it to process long videos and huge documents much faster than its competitors.
  • The Result: On NVIDIA's latest chips (B200), it can produce text output 3 to 9 times faster than similar models while using less memory.

6. What It Can Actually Do (Based on the Paper)

The paper tested this model on specific challenges, and it performed very well in:

  • Reading Documents: It can understand complex charts, tables, and long reports (like financial papers) better than previous versions.
  • Understanding Videos: It can watch a long video with sound and answer questions about what happened, why it happened, and the order of events.
  • Computer Control: It can look at a computer screen (GUI) and figure out which buttons to click to complete a task (like booking a flight or filling out a form).
  • Voice Interaction: It can hold a conversation, understand accents, and answer questions based on what it hears.

7. Open for Everyone

Finally, the paper announces that NVIDIA is sharing the "blueprints" and the "training materials." They are releasing the model in different sizes (standard, compressed, and ultra-compressed) and even sharing some of the data and code they used to build it. This is like giving the recipe and the ingredients to the whole world so other scientists can build their own versions or improve upon them.

In summary: Nemotron 3 Nano Omni is a faster, smarter, multi-sensory assistant that can read, watch, and listen simultaneously, designed to handle long and complex tasks without getting bogged down.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →