Institution profile

Hanhwa Systems

Industry researchasia · kr
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models

Mar 24, 2026

Existing optical flow methods suffer significant performance degradation under realistic image degradations such as motion blur, noise, and compression artifacts. This work proposes a hybrid architecture that integrates intermediate features from diffusion models with convolutional representations, leveraging—for the first time—the inherent degradation-aware capabilities of diffusion models for optical flow estimation. By introducing a cross-frame spatiotemporal attention mechanism, the method enables zero-shot correspondence modeling without requiring retraining or fine-tuning under diverse degradations. The resulting framework establishes a new paradigm for degradation-robust optical flow estimation, consistently outperforming current state-of-the-art approaches across multiple benchmarks under various severe degradation conditions.

0 citationsRead paper

M$^3$KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation

Dec 23, 2025

To address three key challenges in audio-visual multimodal RAG—narrow modality coverage of knowledge graphs, weak multi-hop connectivity, and imprecise retrieval—this paper proposes a query-aligned Multi-hop Multimodal Knowledge Graph (M³KG) construction and retrieval framework. We introduce a novel lightweight multi-agent construction method that significantly expands the modality granularity and cross-modal path depth of multimodal knowledge graphs (MMKGs). Furthermore, we design the GRASP mechanism—comprising query-driven entity anchoring, supportiveness assessment, and redundant context pruning—to enhance retrieval precision. By integrating modality-aware retrieval, query grounding, relevance scoring, and embedding alignment, our approach improves fact consistency and cross-modal localization accuracy for multimodal large language models (MLLMs) in multi-hop reasoning. Extensive evaluation across multiple multimodal benchmarks demonstrates substantial gains in answer faithfulness and reasoning depth.

0 citationsRead paper

Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model

Oct 01, 2025

Traditional recurrent video super-resolution (VSR) methods suffer from gradient vanishing and poor parallelism, while causal Mamba-based models are inherently limited in modeling fine-grained spatial dependencies. To address these issues, we propose an efficient hybrid spatiotemporal modeling architecture. Our approach features: (1) a Gather-Scatter Mamba mechanism that aligns neighboring frame features to the central frame within a temporal window before aggregation and scattering, mitigating occlusion artifacts and enhancing feature redistribution; and (2) integration of shifted-window self-attention to explicitly capture local spatial dependencies, compensating for Mamba’s structural constraints. The architecture retains linear time complexity while enabling precise spatiotemporal feature propagation. Extensive experiments demonstrate state-of-the-art performance on multiple VSR benchmarks, along with significantly accelerated inference—effectively balancing accuracy and efficiency.

0 citationsRead paper
Recent publications

Latest Papers

DA-Flow: Degradation-Aware Optical Flow Estimation with Diffusion Models

Mar 24, 2026

Existing optical flow methods suffer significant performance degradation under realistic image degradations such as motion blur, noise, and compression artifacts. This work proposes a hybrid architecture that integrates intermediate features from diffusion models with convolutional representations, leveraging—for the first time—the inherent degradation-aware capabilities of diffusion models for optical flow estimation. By introducing a cross-frame spatiotemporal attention mechanism, the method enables zero-shot correspondence modeling without requiring retraining or fine-tuning under diverse degradations. The resulting framework establishes a new paradigm for degradation-robust optical flow estimation, consistently outperforming current state-of-the-art approaches across multiple benchmarks under various severe degradation conditions.

0 citationsRead paper

M$^3$KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation

Dec 23, 2025

To address three key challenges in audio-visual multimodal RAG—narrow modality coverage of knowledge graphs, weak multi-hop connectivity, and imprecise retrieval—this paper proposes a query-aligned Multi-hop Multimodal Knowledge Graph (M³KG) construction and retrieval framework. We introduce a novel lightweight multi-agent construction method that significantly expands the modality granularity and cross-modal path depth of multimodal knowledge graphs (MMKGs). Furthermore, we design the GRASP mechanism—comprising query-driven entity anchoring, supportiveness assessment, and redundant context pruning—to enhance retrieval precision. By integrating modality-aware retrieval, query grounding, relevance scoring, and embedding alignment, our approach improves fact consistency and cross-modal localization accuracy for multimodal large language models (MLLMs) in multi-hop reasoning. Extensive evaluation across multiple multimodal benchmarks demonstrates substantial gains in answer faithfulness and reasoning depth.

0 citationsRead paper

Gather-Scatter Mamba: Accelerating Propagation with Efficient State Space Model

Oct 01, 2025

Traditional recurrent video super-resolution (VSR) methods suffer from gradient vanishing and poor parallelism, while causal Mamba-based models are inherently limited in modeling fine-grained spatial dependencies. To address these issues, we propose an efficient hybrid spatiotemporal modeling architecture. Our approach features: (1) a Gather-Scatter Mamba mechanism that aligns neighboring frame features to the central frame within a temporal window before aggregation and scattering, mitigating occlusion artifacts and enhancing feature redistribution; and (2) integration of shifted-window self-attention to explicitly capture local spatial dependencies, compensating for Mamba’s structural constraints. The architecture retains linear time complexity while enabling precise spatiotemporal feature propagation. Extensive experiments demonstrate state-of-the-art performance on multiple VSR benchmarks, along with significantly accelerated inference—effectively balancing accuracy and efficiency.

0 citationsRead paper