英伟达:2025 Nemotron Nano 2技术报告(英文版)(43页).pdf
1、2025-8-18NVIDIA Nemotron Nano 2:An Accurate andEfficient Hybrid Mamba-Transformer ReasoningModelNVIDIAAbstract.We introduce Nemotron-Nano-9B-v2,a hybrid Mamba-Transformer language modeldesigned to increase throughput for reasoning workloads while achieving state-of-the-art accuracycompared to simila
2、rly-sized models.Nemotron-Nano-9B-v2 builds on the Nemotron-H architecture,in which the majority of the self-attention layers in the common Transformer architecture are replacedwith Mamba-2 layers,to achieve improved inference speed when generating the long thinking tracesneeded for reasoning.We cre
3、ate Nemotron-Nano-9B-v2 by first pre-training a 12-billion-parametermodel(Nemotron-Nano-12B-v2-Base)on 20 trillion tokens using an FP8 training recipe.Afteraligning Nemotron-Nano-12B-v2-Base,we employ the Minitron strategy to compress and distillthe model with the goal of enabling inference on up to
4、 128k tokens on a single NVIDIA A10GGPU(22GiB of memory,bfloat16precision).Compared to existing similarly-sized models(e.g.,Qwen3-8B),we show that Nemotron-Nano-9B-v2 achieves on-par or better accuracy on reasoningbenchmarks while achieving up to 6higher inference throughput in reasoning settings li
5、ke 8kinput and 16k output tokens(Figure 1).We are releasing Nemotron-Nano-9B-v2,Nemotron-Nano-12B-v2-Base,and Nemotron-Nano-9B-v2-Base checkpoints along with the majority of our pre-andpost-training datasets on Hugging Face.1.IntroductionWe introduce NVIDIA Nemotron Nano 2,a hybrid Mamba-Transformer
6、 reasoning model(Waleffeet al.,2024;Lieber et al.,2024;DeepMind,2025;NVIDIA,2025)that achieves on-par or betterbenchmark accuracies at 3 6higher throughput than Qwen3-8B(Yang et al.,2025)for generation-heavy scenarios like 1k input/8k output or 8k input/16k output tokens(Figure 1).NemotronNano 2 bui





点击查看更多