Institution profile

National Yang Ming Chiao Tung University

Academic institutionasia · tw
Official website
Research library467linked papers
Opportunities0open roles
Selected work

Representative Papers

DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models

Jul 01, 2024arXiv.org

To address temporal inconsistency arising from direct transfer of pretrained image diffusion inpainting models to video domains, this paper proposes a zero-shot video inpainting framework that reuses arbitrary 2D image diffusion inpainting models without fine-tuning. The method comprises two core innovations: (1) a hierarchical token fusion strategy that enforces inter-frame semantic alignment in the latent space; and (2) a joint optical-flow-guided and feature nearest-neighbor matching mechanism to enhance motion modeling robustness. Crucially, the approach eliminates the need for retraining across diverse degradation types—including 8× super-resolution and Gaussian noise with σ=75—achieving superior performance over fully supervised methods under extreme degradations. Moreover, it demonstrates significant cross-dataset generalization capability, validating its effectiveness beyond domain-specific training.

4 citations2 influentialRead paper

AdaGaR: Adaptive Gabor Representation for Dynamic Scene Reconstruction

Jan 02, 2026arXiv.org

This work addresses the challenge of monocular dynamic 3D scene reconstruction, where existing Gaussian primitive-based methods often suffer from low-pass filtering, energy instability, and interpolation artifacts, leading to a trade-off between high-frequency detail preservation and temporal coherence. To overcome these limitations, we propose AdaGaR, a unified framework that introduces an adaptive Gabor representation with learnable frequency weights and an energy compensation mechanism to enhance high-frequency modeling. Temporal smoothness is ensured through cubic Hermite spline interpolation combined with time-curvature regularization. Furthermore, we design an adaptive initialization strategy that integrates depth estimation, point tracking, and foreground masks. Evaluated on Tap-Vid DAVIS, AdaGaR achieves state-of-the-art performance (PSNR 35.49, SSIM 0.9433, LPIPS 0.0723) and demonstrates strong generalization across tasks including frame interpolation, depth consistency, video editing, and stereo synthesis.

2 citationsRead paper

A Low-Power Streaming Speech Enhancement Accelerator for Edge Devices

Mar 27, 2025IEEE Open Journal of Circuits and Systems

To address the high computational complexity, low energy efficiency, and poor adaptability of Transformer-based speech enhancement models to streaming, low-power edge scenarios, this work proposes a model–hardware co-optimization framework. Methodologically, it introduces domain-aware and streaming-aware joint pruning, a softmax-free attention mechanism, and batch-normalization-enhanced Transformer architecture to improve model lightweighting and hardware friendliness; additionally, it designs a 1D configurable processing array coupled with an SRAM address remapping scheme that eliminates memory skips. Experimental results demonstrate a 93.9% reduction in model size, requiring only 207.8K logic gates and 53.75 KB SRAM. Operating at 62.5 MHz, the system achieves a mere 8.08 mW power consumption while enabling real-time streaming speech denoising—significantly outperforming state-of-the-art edge speech enhancement solutions.

1 citations1 influentialRead paper

StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusions

Oct 02, 2025

This work exposes the security vulnerability of 3D Gaussian Splatting (3DGS) to image-level poisoning attacks and proposes the first density-guided, stealthy poisoning method. To address the problem of maintaining high-fidelity reconstruction in benign views while inducing hallucinated objects in targeted views, the method leverages kernel density estimation (KDE) to identify sparse regions in the scene and injects view-dependent Gaussian primitives therein. It further incorporates adaptive noise to disrupt multi-view consistency, enhancing both stealthiness and robustness against detection and defense. A KDE-based evaluation protocol is introduced for objective, quantitative assessment of poisoning efficacy. Experiments demonstrate that the proposed approach achieves significantly higher attack success rates than existing state-of-the-art methods, while preserving reconstruction quality in untargeted (innocent) views. The method exhibits strong visual deception capability and resilience against common defenses.

1 citationsRead paper

Traffic Scene Generation from Natural Language Description for Autonomous Vehicles with Large Language Model

Sep 15, 2024arXiv.org

To address insufficient diversity and limited coverage of critical scenarios in natural language–driven traffic scene generation for autonomous driving simulation, this paper proposes the first large language model (LLM)–driven end-to-end text-to-scene generation framework. The method integrates semantic parsing, vector-based retrieval, multi-factor road ranking, and joint planning of dynamic road networks and agent behaviors—thereby overcoming reliance on predefined trajectories. It enables semantically controllable generation of both routine and high-risk driving scenarios and seamlessly interfaces with the CARLA simulator. Evaluated on the SafeBench benchmark, the framework reduces the average collision rate from 8.0% to 3.5%, while significantly improving narrative coherence and causal reasoning in scene descriptions. This work establishes a scalable, interpretable paradigm for safety-critical scenario generation in autonomous driving validation.

1 citationsRead paper
Recent publications

Latest Papers