NVIDIA Nemotron Nano V2 VL

초록

우리는 강력한 실제 문서 이해, 장편 비디오 이해 및 추론 과제를 위해 설계된 Nemotron 비전-언어 시리즈의 최신 모델인 Nemotron Nano V2 VL을 소개합니다. Nemotron Nano V2 VL은 모델 아키텍처, 데이터셋 및 학습 방법론의 주요 개선을 통해 모든 비전 및 텍스트 영역에서 이전 모델인 Llama-3.1-Nemotron-Nano-VL-8B 대비 상당한 향상을 제공합니다. Nemotron Nano V2 VL은 하이브리드 Mamba-Transformer LLM인 Nemotron Nano V2와 혁신적인 토큰 축소 기술을 기반으로 하여 장문 문서 및 비디오 시나리오에서 더 높은 추론 처리량을 달성합니다. BF16, FP8 및 FP4 형식의 모델 체크포인트와 데이터셋, 방법론 및 학습 코드의 상당 부분을 공개합니다.

English

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant improvements over our previous model, Llama-3.1-Nemotron-Nano-VL-8B, across all vision and text domains through major enhancements in model architecture, datasets, and training recipes. Nemotron Nano V2 VL builds on Nemotron Nano V2, a hybrid Mamba-Transformer LLM, and innovative token reduction techniques to achieve higher inference throughput in long document and video scenarios. We are releasing model checkpoints in BF16, FP8, and FP4 formats and sharing large parts of our datasets, recipes and training code.