Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过混合分析-学习方法研究了流匹配模型中生成样本的归因问题,提出了基于轨迹的归因评分,并在不同基准上进行了评估。
📝 Abstract
Understanding which training samples influence a generated image is an important problem in generative modeling. In flow matching, training samples influence the generated image through the velocity field along the generation trajectory. Removing samples to examine their counterfactual influence changes the velocity field, and the resulting effect on the final image depends on how the change propagates through the trajectory. Consequently, local changes in the velocity field do not necessarily predict the final counterfactual effect. This work investigates attribution in flow-matching models through a hybrid analytical--learned approach, and uses it to derive trajectory-based attribution scores at the cluster level. We evaluate these attribution scores using independently retrained leave-one-cluster-out (LOO) models, and compare with several attribution baselines using two different flow-matching latent spaces. Our experiments show that semantic similarity constitutes a strong baseline, while the closed-form trajectory-based attribution is competitive in some metrics without requiring counterfactual retraining or model gradients. Our results show that attribution in flow matching depends not only on semantic similarity to training samples, but also on the latent representation, trajectory dynamics, and how influence is propagated to the final output.
Problem

Research questions and friction points this paper is trying to address.

generated image
training samples
flow matching
velocity field
generation trajectory
Innovation

Methods, ideas, or system contributions that make the work stand out.

flow matching
attribution
trajectory-based attribution
latent representation
semantic similarity