Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过将提示矩阵分解为低秩形式,减少了视觉-语言模型中提示学习的参数量,并发现固定token侧因子对性能影响不大,从而在少量样本下实现更高效的学习。
📝 Abstract
Prompt learning adapts CLIP to downstream recognition by replacing hand-written templates with learned continuous context vectors, which in Context Optimization (CoOp) form a dense prompt matrix $\mathbf{P}\in\mathbb{R}^{m\times d}$ trained from only a few examples per class. We study whether this matrix is over-parameterized by factorizing it as $\mathbf{P}=\mathbf{B}\mathbf{A}$, which cuts the trainable prompt parameters from $md$ to $r(m+d)$, and to $rd$ once the token-side factor $\mathbf{B}$ is fixed. Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization. We then find that the token-side factor need not be learned at all: fixing $\mathbf{B}$ to a Gaussian, orthogonal, SVD-derived, or even random basis and training only the embedding-side factor $\mathbf{A}$ stays on par with the fully trainable factorization, and a source-trained $\mathbf{B}$ offers no advantage over a random one. A prompt-factor asymmetry and a local update-space dimension gap show why fixing $\mathbf{B}$ is far less restrictive than fixing $\mathbf{A}$, and a smoothness-only guarantee certifies that optimizing $\mathbf{A}$ over a fixed $\mathbf{B}$ converges. In the CLIP prompt setting, the embedding-side coefficients carry the adaptation while the token basis can simply be fixed.
Problem

Research questions and friction points this paper is trying to address.

Low-Rank Prompt Learning
Vision-Language Models
Fixed-Token Bases
Parameter Reduction
Few-Shot Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-Rank Prompt Learning
Fixed-Token Bases
Parameter Efficiency
Few-Shot Generalization
🔎 Similar Papers
No similar papers found.