Institution profile

Nomic AI

Industry researchnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

NOMAD Projection

May 21, 2025

Generative AI has triggered an explosion in data volume, rendering traditional nonlinear dimensionality reduction methods—such as t-SNE and UMAP—ineffective for scaling to million-scale unstructured embeddings, thereby severely hindering exploratory data analysis in AI interpretability. Method: We propose the first scalable visualization framework supporting multi-GPU distributed training. It introduces an information-theoretic upper bound approximation of the InfoNC-t-SNE loss, integrated with deep metric learning that combines negative sampling and mean-affinity discrimination. Contribution/Results: Our framework achieves end-to-end mapping of the full Multilingual Wikipedia embedding corpus (>10 million entries). Experiments demonstrate substantial improvements over state-of-the-art methods in both speed and visualization quality. Notably, it produces the first global semantic map for multi-lingual text embeddings at the ten-million scale, establishing a novel paradigm for large-scale AI interpretability.

0 citationsRead paper

Training Sparse Mixture Of Experts Text Embedding Models

Feb 11, 2025

To address the deployment bottlenecks—namely high inference latency and excessive memory consumption—posed by large-parameter Transformer-based text embedding models in retrieval-augmented generation (RAG) scenarios, this work pioneers the integration of sparse Mixture-of-Experts (MoE) architecture into general-purpose text embedding. We propose Nomic Embed v2, the first open-source, general-purpose sparse MoE text embedding model. Built upon a Transformer encoder, it employs gated routing and multi-task joint training—including contrastive learning—to achieve substantial reductions in memory footprint and latency at equivalent parameter counts. Empirically, Nomic Embed v2 surpasses comparable models across diverse monolingual and multilingual benchmarks (e.g., MIRACL, BEIR), demonstrating superior cross-lingual robustness. Notably, its performance matches that of dense models twice its size. All code, model weights, and evaluation datasets are fully open-sourced.

0 citationsRead paper
Recent publications

Latest Papers

NOMAD Projection

May 21, 2025

Generative AI has triggered an explosion in data volume, rendering traditional nonlinear dimensionality reduction methods—such as t-SNE and UMAP—ineffective for scaling to million-scale unstructured embeddings, thereby severely hindering exploratory data analysis in AI interpretability. Method: We propose the first scalable visualization framework supporting multi-GPU distributed training. It introduces an information-theoretic upper bound approximation of the InfoNC-t-SNE loss, integrated with deep metric learning that combines negative sampling and mean-affinity discrimination. Contribution/Results: Our framework achieves end-to-end mapping of the full Multilingual Wikipedia embedding corpus (>10 million entries). Experiments demonstrate substantial improvements over state-of-the-art methods in both speed and visualization quality. Notably, it produces the first global semantic map for multi-lingual text embeddings at the ten-million scale, establishing a novel paradigm for large-scale AI interpretability.

0 citationsRead paper

Training Sparse Mixture Of Experts Text Embedding Models

Feb 11, 2025

To address the deployment bottlenecks—namely high inference latency and excessive memory consumption—posed by large-parameter Transformer-based text embedding models in retrieval-augmented generation (RAG) scenarios, this work pioneers the integration of sparse Mixture-of-Experts (MoE) architecture into general-purpose text embedding. We propose Nomic Embed v2, the first open-source, general-purpose sparse MoE text embedding model. Built upon a Transformer encoder, it employs gated routing and multi-task joint training—including contrastive learning—to achieve substantial reductions in memory footprint and latency at equivalent parameter counts. Empirically, Nomic Embed v2 surpasses comparable models across diverse monolingual and multilingual benchmarks (e.g., MIRACL, BEIR), demonstrating superior cross-lingual robustness. Notably, its performance matches that of dense models twice its size. All code, model weights, and evaluation datasets are fully open-sourced.

0 citationsRead paper