🤖 AI Summary
This study addresses the lack of a systematic evaluation framework for underwater image reconstruction, which has hindered comprehensive assessment of methods in terms of accuracy, viewpoint consistency, and robustness to varying water conditions. To bridge this gap, the authors propose the first multidimensional evaluation framework that jointly considers reconstruction fidelity, camera motion consistency, and the impact of water quality, accompanied by a newly curated real-world underwater image dataset for benchmarking. Through extensive experiments comparing traditional physics-based scattering models with emerging vision-language models (VLMs), the results demonstrate that VLMs—despite eschewing explicit physical modeling—consistently outperform conventional approaches across all evaluated metrics, achieving superior reconstruction quality and generalization capability.
📝 Abstract
Underwater image restoration consists of recovering an image which looks like there is no water present. To date, evaluation has not been systematic. This paper describes a systematic evaluation pipeline for underwater reconstruction, which can be used to assess a method for accuracy; consistency of reconstruction over camera moves; and the effect of water parameters. We use this pipeline to evaluate a range of current procedures, from models constructed using explicit but approximate physical models of scattering to Vision-Language Models (VLMs which are not currently trained with explicit physical models). Overall, VLMs wholly and significantly outperform physically based models in our evaluation, likely because of the importance of a strong image prior. Results on images of real underwater scenes strongly confirm the evaluation.