Distance-to-Distance Ratio: A Similarity Measure for Sentences Based on Rate of Change in LLM Embeddings

📅 2026-01-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current text embedding similarity metrics struggle to accurately capture human perception of subtle semantic differences between sentences. This work proposes Distance Ratio (DDR) as a novel similarity measure for sentence embeddings derived from large language models. DDR quantifies the relative rate of change in word embedding similarity before and after contextualization, thereby characterizing how context modulates semantic interpretation. Inspired by Lipschitz continuity, DDR introduces—for the first time—the rate of similarity change in embedding space as a semantic discriminant, significantly enhancing sensitivity to fine-grained semantic distinctions. Experimental results demonstrate that under controlled textual edits, DDR more precisely differentiates between semantically similar and dissimilar sentence variants compared to prevailing similarity metrics.

Technology Category

Application Category

📝 Abstract
A measure of similarity between text embeddings can be considered adequate only if it adheres to the human perception of similarity between texts. In this paper, we introduce the distance-to-distance ratio (DDR), a novel measure of similarity between LLM sentence embeddings. Inspired by Lipschitz continuity, DDR measures the rate of change in similarity between the pre-context word embeddings and the similarity between post-context LLM embeddings, thus measuring the semantic influence of context. We evaluate the performance of DDR in experiments designed as a series of perturbations applied to sentences drawn from a sentence dataset. For each sentence, we generate variants by replacing one, two, or three words with either synonyms, which constitute semantically similar text, or randomly chosen words, which constitute semantically dissimilar text. We compare the performance of DDR with other prevailing similarity metrics and demonstrate that DDR consistently provides finer discrimination between semantically similar and dissimilar texts, even under minimal, controlled edits.
Problem

Research questions and friction points this paper is trying to address.

sentence similarity
LLM embeddings
semantic similarity
context influence
similarity measure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distance-to-Distance Ratio
LLM embeddings
semantic similarity
contextual influence
Lipschitz continuity
A
Abdullah Qureshi
DataKnife, Chicago IL 60611, USA
Kenneth Rice
Kenneth Rice
University of Washington
Biostatistics
A
Alexander Wolpert
Roosevelt University, Chicago IL 60605, USA