Institution profile

Nanjing University

Academic institutionasia · cn
Official website
Research library2,325linked papers
Opportunities0open roles
Selected work

Representative Papers

A Survey of Deep Face Restoration: Denoise, Super-Resolution, Deblur, Artifact Removal

Nov 05, 2022arXiv.org

This paper presents a systematic survey of deep learning–based facial image restoration, focusing on denoising, super-resolution, deblurring, and artifact removal. Addressing challenges such as strong facial structural priors and complex degradation modeling, we propose the first holistic taxonomy of methods, a unified evaluation framework, and an open-source benchmark repository encompassing 20+ state-of-the-art approaches—including fully reproducible implementations. Leveraging datasets like CelebA and FFHQ, we conduct comprehensive cross-method evaluations using PSNR, SSIM, and LPIPS metrics, integrating CNN/Transformer architectures, perceptual and adversarial losses, and multi-scale feature fusion strategies. Our empirical analysis reveals performance boundaries and task-specific suitability across methods. Key contributions include: (1) the first structured, principle-driven classification system for facial restoration; (2) the first open, end-to-end benchmark platform supporting full method reproduction; (3) a rigorous, large-scale empirical study; and (4) concrete research directions concerning network design, evaluation paradigms, and dataset construction.

40 citations1 influentialRead paper

A Systematic Literature Review on Large Language Models for Automated Program Repair

May 02, 2024arXiv.org

Research on large language models (LLMs) for automated program repair (APR) remains fragmented and lacks a systematic, unified understanding. Method: We conduct a systematic literature review (SLR) covering 127 papers published between 2020 and 2024, establishing the first comprehensive conceptual framework for LLM-based APR. We categorize model utilization strategies into three types—fine-tuning, prompt engineering, and hybrid ensemble—and perform multidimensional thematic analysis across input representation, semantic/security-specific repair scenarios, and open-science practices. Contribution/Results: We identify core challenges including model robustness, evaluation bias, and real-world deployment adaptability. The study yields a reusable taxonomy, benchmark insights, and methodological guidelines—delivering the APR community’s first holistic landscape map to precisely identify research gaps and inform future innovation pathways.

39 citations1 influentialRead paper

Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation

Dec 17, 2023

This work addresses the semantic gap between natural language instructions and robotic physical actions to enhance the naturalness and reliability of human-robot collaboration. We propose the first four-dimensional taxonomy for language-conditioned robotic manipulation—comprising reward shaping, policy learning, neurosymbolic AI, and foundation model–driven approaches—and systematically analyze their fundamental limitations in generalization and safety. Integrating large language models (LLMs), vision-language models (VLMs), neurosymbolic reasoning, and multimodal semantic parsing, we develop a unified analytical framework spanning semantic extraction, environmental assessment, and auxiliary task design. Our analysis rigorously characterizes the performance boundaries of each paradigm for the first time, establishing theoretical foundations and concrete technical pathways toward safe, generalizable, and interpretable language-driven robotic systems.

10 citationsRead paper

Revisiting Weighted Strategy for Non-stationary Parametric Bandits

Mar 05, 2023International Conference on Artificial Intelligence and Statistics

This work addresses the limitations of existing approaches to non-stationary parametric bandits and Markov decision processes (MDPs), where inadequate analysis of weighted policies often leads to algorithmic complexity and suboptimal efficiency. To overcome this, the paper introduces a refined analytical framework for weighted policy optimization that integrates weighted regression, function approximation, and dynamic regret analysis. This framework substantially simplifies algorithm design while achieving tighter dynamic regret bounds across multiple settings: it matches the performance of sliding-window or restart-based methods in linear bandits, improves the regret bound in generalized linear bandits (GLBs) from ~O(T^{4/5}) to ~O(T^{3/4}), and provides the first dynamic regret guarantees based on weighted policies for two classes of non-stationary MDPs.

9 citations4 influentialRead paper

OmniBench: Towards The Future of Universal Omni-Language Models

Sep 23, 2024arXiv.org

Existing open-source multimodal large language models (MLLMs) exhibit significant deficiencies in joint visual-auditory-textual understanding and reasoning, achieving only ~50% instruction-following accuracy on trilingual multimodal tasks. Method: We introduce OmniBench—the first benchmark for trilingual multimodal collaborative reasoning—and formalize the omni-language model (OLM), a unified architecture capable of jointly processing visual, auditory, and textual (V-A-T) inputs. We construct OmniBench via expert human annotation across diverse trilingual multimodal tasks and curate OmniInstruct, a large-scale instruction-tuning dataset comprising 96K samples. Our methodology integrates cross-modal alignment modeling, trilingual multimodal instruction tuning, and a human-in-the-loop evaluation framework. Contribution/Results: Experiments reveal severe generalization limitations of current open-source OLMs on trilingual multimodal tasks; OmniInstruct substantially improves their reasoning performance. This work establishes a novel evaluation paradigm, provides high-quality resources, and outlines a technical pathway for advancing trilingual multimodal foundation models.

9 citations2 influentialRead paper
Recent publications

Latest Papers