NOMAD Projection
Generative AI has triggered an explosion in data volume, rendering traditional nonlinear dimensionality reduction methods—such as t-SNE and UMAP—ineffective for scaling to million-scale unstructured embeddings, thereby severely hindering exploratory data analysis in AI interpretability. Method: We propose the first scalable visualization framework supporting multi-GPU distributed training. It introduces an information-theoretic upper bound approximation of the InfoNC-t-SNE loss, integrated with deep metric learning that combines negative sampling and mean-affinity discrimination. Contribution/Results: Our framework achieves end-to-end mapping of the full Multilingual Wikipedia embedding corpus (>10 million entries). Experiments demonstrate substantial improvements over state-of-the-art methods in both speed and visualization quality. Notably, it produces the first global semantic map for multi-lingual text embeddings at the ten-million scale, establishing a novel paradigm for large-scale AI interpretability.