Institution profile

Ningbo University of Technology

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

An AI4AI Framework for Visual Token Pruning

Aug 07, 2026

This work addresses the limitations of existing vision token pruning methods, which rely on handcrafted heuristics and struggle to generalize across diverse model architectures and pruning objectives. To overcome this, the authors propose AutoPrune, a training-free AI4AI framework that leverages large language models (LLMs) to automatically design effective pruning strategies. Key innovations include TPDSL, a domain-specific language tailored for token pruning; a residual representation of search states to emphasize critical strategy components; and structured guidance for LLMs to generate algorithms within a constrained space. Extensive experiments demonstrate AutoPrune’s superior performance across 14 benchmarks and three multimodal large language models: it retains over 99% of original accuracy even after removing 94.4% of visual tokens, achieving a 9.9× reduction in FLOPs and a 6.4× speedup in prefill latency.

0 citationsRead paper

DINVMark: A Deep Invertible Network for Video Watermarking

Sep 22, 2025

Existing video watermarking methods suffer from limited embedding capacity, insufficient robustness against HEVC compression, and lack of end-to-end differentiability. To address these limitations, this paper proposes a deep invertible neural network (INN)-based video watermarking framework. Its key contributions are: (1) a differentiable HEVC compression simulation layer that accurately models real-world encoding distortions; (2) a shared encoder-decoder invertible architecture enabling tightly coupled, fully reversible watermark embedding and extraction; and (3) end-to-end joint optimization balancing visual fidelity, embedding capacity, and compression robustness. Experiments demonstrate that, under comparable BD-Rate, the method achieves 2.3× higher watermark capacity, maintains >98.5% extraction accuracy under HEVC compression (CRF=22–37), and yields reconstructed video PSNR >38 dB—significantly outperforming state-of-the-art approaches.

0 citationsRead paper
Recent publications

Latest Papers

An AI4AI Framework for Visual Token Pruning

Aug 07, 2026

This work addresses the limitations of existing vision token pruning methods, which rely on handcrafted heuristics and struggle to generalize across diverse model architectures and pruning objectives. To overcome this, the authors propose AutoPrune, a training-free AI4AI framework that leverages large language models (LLMs) to automatically design effective pruning strategies. Key innovations include TPDSL, a domain-specific language tailored for token pruning; a residual representation of search states to emphasize critical strategy components; and structured guidance for LLMs to generate algorithms within a constrained space. Extensive experiments demonstrate AutoPrune’s superior performance across 14 benchmarks and three multimodal large language models: it retains over 99% of original accuracy even after removing 94.4% of visual tokens, achieving a 9.9× reduction in FLOPs and a 6.4× speedup in prefill latency.

0 citationsRead paper

DINVMark: A Deep Invertible Network for Video Watermarking

Sep 22, 2025

Existing video watermarking methods suffer from limited embedding capacity, insufficient robustness against HEVC compression, and lack of end-to-end differentiability. To address these limitations, this paper proposes a deep invertible neural network (INN)-based video watermarking framework. Its key contributions are: (1) a differentiable HEVC compression simulation layer that accurately models real-world encoding distortions; (2) a shared encoder-decoder invertible architecture enabling tightly coupled, fully reversible watermark embedding and extraction; and (3) end-to-end joint optimization balancing visual fidelity, embedding capacity, and compression robustness. Experiments demonstrate that, under comparable BD-Rate, the method achieves 2.3× higher watermark capacity, maintains >98.5% extraction accuracy under HEVC compression (CRF=22–37), and yields reconstructed video PSNR >38 dB—significantly outperforming state-of-the-art approaches.

0 citationsRead paper