NVIDIA released Nemotron 3 Embed on July 16, 2026, a collection of three open, commercially usable embedding models: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16 and a Blackwell-optimized Nemotron-3-Embed-1B-NVFP4. All three have a 32k context window, mean pooling and query/document prefixes; the 8B model outputs 4096-dimension embeddings and the 1.14B variants 2048. NVIDIA reports the 8B model ranked #1 on the RTEB multilingual leaderboard as of July 15, 2026, scoring 78.5% on RTEB and 75.5% on MMTEB Retrieval, while the 1B BF16 scored 72.4% and 71.0%. The 8B adapts a Ministral-3-8B-Instruct-2512 backbone into a bidirectional encoder; the 1B models come from pruning and distilling a 3B retriever with NVIDIA ModelOpt. Weights, datasets, fine-tuning and distillation recipes are on Hugging Face, with vLLM support and an NVIDIA NIM microservice for the 1B; Elastic, turbopuffer, Mem0, Zep, Boomi, IBM, Palantir, ServiceNow and Zoom are integrating or evaluating the models.
Sources