Institution profile

Shenzhen University of Advanced Technology

Academic institutionasia · cn
Official website
Research library212linked papers
Opportunities0open roles
Selected work

Representative Papers

MagicFight: Personalized Martial Arts Combat Video Generation

Oct 28, 2024ACM Multimedia

This work addresses the challenges of identity confusion, anatomical implausibility, and motion incoherence commonly observed in existing personalized video generation methods when applied to two-person martial arts sparring scenarios. To tackle this, we introduce and implement the first personalized dual-character martial arts combat video generation task, thereby filling a critical gap in interactive human video synthesis for complex dyadic interactions. We construct a high-quality 3D sparring dataset using the Unity physics engine and propose a tailored generative model that integrates identity-preserving mechanisms with motion coordination constraints. Experimental results demonstrate that our approach produces high-fidelity videos featuring consistent character identities, temporally coherent movements, and realistic interaction dynamics, establishing a new paradigm for interactive content creation.

23 citationsRead paper

Seal2Real: Prompt Prior Learning on Diffusion Model for Unsupervised Document Seal Data Generation and Realisation

Oct 01, 2023arXiv.org

Document seal processing tasks—including segmentation, authenticity verification, removal, and occluded text recognition—are severely hindered by the scarcity of real-world annotated data. To address this, we propose the first end-to-end unsupervised seal image generation framework, built upon Stable Diffusion. Our method introduces a novel prompt-based prior learning mechanism that enables structural controllability and high-fidelity synthesis without requiring paired real seal samples. Leveraging this framework, we construct Seal-DB—the first large-scale, fully annotated seal dataset comprising 20,000 images. Extensive evaluation on Seal-DB demonstrates significant improvements in downstream tasks: seal segmentation and occluded text recognition accuracy increase by 12.6%–18.3%. Moreover, expert blind evaluation confirms that 91.4% of generated seals are perceived as photorealistic. This work establishes a foundational resource and methodology for data-starved seal analysis research.

4 citationsRead paper

Efficient Chest X-ray Representation Learning via Semantic-Partitioned Contrastive Learning

Mar 07, 2026

Existing self-supervised methods for chest X-ray (CXR) analysis are often limited by their reliance on strong data augmentations, sensitivity to non-diagnostic background regions, or high computational costs. This work proposes Semantic-aware Patch-based Contrastive Learning (S-PCL), which constructs an intrinsic information bottleneck by randomly partitioning image patches from a single CXR into two complementary yet incomplete semantic subsets. This design compels the encoder to infer global anatomical and pathological structures from local cues, thereby modeling long-range dependencies and structural consistency. Notably, S-PCL operates without handcrafted augmentations, momentum encoders, or auxiliary decoders. It achieves state-of-the-art accuracy on benchmark datasets such as ChestX-ray14 and CheXpert while significantly reducing computational overhead—requiring the lowest GFLOPs among current self-supervised approaches.

1 citationsRead paper

The Model Knows Which Tokens Matter: Automatic Token Selection via Noise Gating

Mar 07, 2026

This work addresses the redundancy introduced by excessive visual tokens in vision-language models, which significantly increases inference overhead. The authors formulate token pruning as a bandwidth-constrained information transmission problem and propose an efficient, annotation-free pruning method that requires no auxiliary objectives. By employing lightweight Scorer and Denoiser modules, the approach learns to predict token importance using only the standard language modeling loss. A variance-preserving noisy gating mechanism ensures full gradient flow during training while enabling hard top-K selection at inference time. The method is architecture-agnostic and achieves strong transferability, retaining 96.5% of the original accuracy across ten vision-language benchmarks, accelerating LLM prefilling by 2.85×, and adding merely 0.69 milliseconds of latency.

1 citationsRead paper

LegalMALR:Multi-Agent Query Understanding and LLM-Based Reranking for Chinese Statute Retrieval

Jan 25, 2026

This study addresses the challenge that real-world legal queries often involve multiple intertwined issues and are expressed in informal or ambiguous language, which hinders traditional retrieval methods from accurately recalling relevant statutes. To overcome this limitation, the authors propose a novel paradigm integrating multi-agent query understanding with zero-shot large language model–based re-ranking. The approach leverages multi-perspective query reformulation, iterative dense retrieval guided by Generalized Reinforcement Policy Optimization (GRPO), and natural language legal reasoning to achieve precise statute localization. Evaluated on both the newly constructed CSAID dataset and the public STARD benchmark, the method significantly outperforms existing retrieval-augmented generation (RAG) approaches, demonstrating superior performance in both in-distribution and out-of-distribution scenarios.

1 citationsRead paper
Recent publications

Latest Papers