Zero2Text: Zero-Training Cross-Domain Inversion Attacks on Textual Embeddings
This work addresses the challenge of embedding inversion under strict black-box and cross-domain settings, where existing methods struggle to preserve vector database privacy due to their reliance on extensive queries or in-domain training data. We propose the first training-free cross-domain embedding inversion framework, which leverages a recursive online alignment mechanism that integrates large language model priors with dynamic ridge regression to generate text aligned with target embeddings in real time—without any training or prior knowledge of the target domain. Our approach eliminates dependence on static datasets or high query budgets and exposes limitations of conventional defenses such as differential privacy. Experiments demonstrate significant improvements over baselines across multiple benchmarks including MS MARCO, achieving a 1.8× increase in ROUGE-L and a 6.4× gain in BLEU-2 against OpenAI models, successfully reconstructing original sentences from unseen domains.