HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering
This work addresses the challenge of continual visual question answering, where shared parameter updates often lead to interference across tasks and objects, hindering the simultaneous adaptation to new tasks and retention of old knowledge under non-stationary data streams. To mitigate this, the authors propose a novel paradigm that constructs a drift-resistant memory bank of anchors encoding both visual objects and textual task representations. A hypernetwork dynamically generates lightweight, low-rank adapters (LoRAs) conditioned on retrieved anchors to enable precise and efficient adaptation to the current task. Additionally, a semantic–parameter space alignment loss is introduced to reduce interference and enhance knowledge stability. The proposed method significantly outperforms state-of-the-art approaches on both VQA v2 and NExT-QA benchmarks under standard and compositional continual learning settings.