Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
研究提出EyeControl,一种基于MLLM的图像润饰代理,通过弱用户意图指导视觉焦点增强,实现自然协调的局部和全局调整。
研究提出EyeControl,一种基于MLLM的图像润饰代理,通过弱用户意图指导视觉焦点增强,实现自然协调的局部和全局调整。
Traditional single-hypothesis optical flow often leads to ghosting, structural distortion, or blurriness in interpolated frames when confronted with ambiguous matching regions such as repetitive textures, symmetric structures, or motion blur. This work proposes the first multi-hypothesis optical flow estimation framework that operates without ground-truth flow supervision. By maintaining multiple candidate correspondences and employing a reliability-guided routing mechanism to select the optimal hypothesis, the method avoids soft blending and instead refines each hypothesis independently through anchor initialization and local attention. This approach significantly enhances interpolation accuracy, achieving state-of-the-art performance in terms of LPIPS and DISTS metrics on MA-HD and several standard video frame interpolation benchmarks, effectively mitigating ghosting artifacts and structural distortions.
This work addresses the ghosting artifacts commonly encountered in traditional multi-exposure HDR imaging of dynamic scenes, which stem from the use of fixed feedforward reconstruction paradigms. To overcome this limitation, the paper introduces an adaptive reconstruction framework that incorporates an agent-based mechanism into HDR imaging for the first time. The proposed approach leverages a multimodal large language model to enable scene understanding and dynamic region parsing, and integrates fine-grained contextual knowledge matching, a perception–distortion feedback loop, and an agent-guided generative alignment strategy to dynamically optimize the reconstruction process. This method significantly suppresses ghosting and local artifacts, achieving state-of-the-art or superior performance in both objective metrics and visual quality.
This work addresses the limited generalization of existing millimeter-wave radar-based human activity recognition methods across heterogeneous radar sources—such as varying devices and frequency bands—by introducing UniMM-HAR, the first large-scale heterogeneous multi-source mmWave point cloud dataset for activity recognition, along with a novel Doppler-aware point cloud network, DAP-Net. DAP-Net incorporates a Dual-space Doppler Re-parameterization (D2R) module and a Text-aligned Anchor Mechanism (TAM) that leverages semantic text alignment to stabilize feature learning. Through cross-modal feature alignment and Doppler-pattern-guided feature recalibration, the model enhances robustness under distribution shifts. Experimental results demonstrate that DAP-Net significantly outperforms current state-of-the-art methods in heterogeneous settings, achieving new benchmark accuracy and exhibiting strong cross-source generalization capabilities.
This work addresses the limitation of existing 2D-to-3D conversion methods, which, despite achieving geometric accuracy, often lack artistic expressiveness and fail to replicate the immersive depth and emotional resonance characteristic of professional 3D cinema. To bridge this gap, the authors propose a novel paradigm termed “artistic disparity synthesis,” implemented through the Art3D framework. This approach decouples global depth parameters from local artistic effects, shifting the conversion objective from physically accurate disparity estimation to artistically consistent disparity generation. Leveraging a dual-path architecture and indirect supervision trained on professional 3D film data, the method enables art-directed depth modeling. A new quantitative metric is introduced to evaluate alignment with cinematic style. Experiments demonstrate that the proposed method successfully reproduces key out-of-screen pop-out effects and achieves high alignment with the global depth aesthetics of professional 3D content, thereby validating the feasibility of art-driven stereoscopic conversion.
研究提出EyeControl,一种基于MLLM的图像润饰代理,通过弱用户意图指导视觉焦点增强,实现自然协调的局部和全局调整。
Traditional single-hypothesis optical flow often leads to ghosting, structural distortion, or blurriness in interpolated frames when confronted with ambiguous matching regions such as repetitive textures, symmetric structures, or motion blur. This work proposes the first multi-hypothesis optical flow estimation framework that operates without ground-truth flow supervision. By maintaining multiple candidate correspondences and employing a reliability-guided routing mechanism to select the optimal hypothesis, the method avoids soft blending and instead refines each hypothesis independently through anchor initialization and local attention. This approach significantly enhances interpolation accuracy, achieving state-of-the-art performance in terms of LPIPS and DISTS metrics on MA-HD and several standard video frame interpolation benchmarks, effectively mitigating ghosting artifacts and structural distortions.
This work addresses the ghosting artifacts commonly encountered in traditional multi-exposure HDR imaging of dynamic scenes, which stem from the use of fixed feedforward reconstruction paradigms. To overcome this limitation, the paper introduces an adaptive reconstruction framework that incorporates an agent-based mechanism into HDR imaging for the first time. The proposed approach leverages a multimodal large language model to enable scene understanding and dynamic region parsing, and integrates fine-grained contextual knowledge matching, a perception–distortion feedback loop, and an agent-guided generative alignment strategy to dynamically optimize the reconstruction process. This method significantly suppresses ghosting and local artifacts, achieving state-of-the-art or superior performance in both objective metrics and visual quality.
This work addresses the limited generalization of existing millimeter-wave radar-based human activity recognition methods across heterogeneous radar sources—such as varying devices and frequency bands—by introducing UniMM-HAR, the first large-scale heterogeneous multi-source mmWave point cloud dataset for activity recognition, along with a novel Doppler-aware point cloud network, DAP-Net. DAP-Net incorporates a Dual-space Doppler Re-parameterization (D2R) module and a Text-aligned Anchor Mechanism (TAM) that leverages semantic text alignment to stabilize feature learning. Through cross-modal feature alignment and Doppler-pattern-guided feature recalibration, the model enhances robustness under distribution shifts. Experimental results demonstrate that DAP-Net significantly outperforms current state-of-the-art methods in heterogeneous settings, achieving new benchmark accuracy and exhibiting strong cross-source generalization capabilities.
This work addresses the limitation of existing 2D-to-3D conversion methods, which, despite achieving geometric accuracy, often lack artistic expressiveness and fail to replicate the immersive depth and emotional resonance characteristic of professional 3D cinema. To bridge this gap, the authors propose a novel paradigm termed “artistic disparity synthesis,” implemented through the Art3D framework. This approach decouples global depth parameters from local artistic effects, shifting the conversion objective from physically accurate disparity estimation to artistically consistent disparity generation. Leveraging a dual-path architecture and indirect supervision trained on professional 3D film data, the method enables art-directed depth modeling. A new quantitative metric is introduced to evaluate alignment with cinematic style. Experiments demonstrate that the proposed method successfully reproduces key out-of-screen pop-out effects and achieves high alignment with the global depth aesthetics of professional 3D content, thereby validating the feasibility of art-driven stereoscopic conversion.