Release2026-07-16

NVIDIA released Nemotron 3 Embed on July 16, 2026, a collection of three open, commercially usable embedding models: Nemotron-3-Embed-8B-BF16, Nemotron-3-Embed-1B-BF16 and a Blackwell-optimized Nemotron-3-Embed-1B-NVFP4. All three have a 32k context window, mean pooling and query/document prefixes; the 8B model outputs 4096-dimension embeddings and the 1.14B variants 2048. NVIDIA reports the 8B model ranked #1 on the RTEB multilingual leaderboard as of July 15, 2026, scoring 78.5% on RTEB and 75.5% on MMTEB Retrieval, while the 1B BF16 scored 72.4% and 71.0%. The 8B adapts a Ministral-3-8B-Instruct-2512 backbone into a bidirectional encoder; the 1B models come from pruning and distilling a 3B retriever with NVIDIA ModelOpt. Weights, datasets, fine-tuning and distillation recipes are on Hugging Face, with vLLM support and an NVIDIA NIM microservice for the 1B; Elastic, turbopuffer, Mem0, Zep, Boomi, IBM, Palantir, ServiceNow and Zoom are integrating or evaluating the models.

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.