Institution profile

TCL Communication

Industry researchasia · cn
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

Lifted Gabidulin Construction for LDPC Representations of Finite Geometry Codes

Jun 09, 2026

This work addresses the poor iterative decoding performance of low-density parity-check (LDPC) codes derived from finite geometries, which stems from the dense structure and abundance of short cycles in their conventional parity-check matrices. To overcome this limitation, the authors propose a sparsification method based on pencil selection, formulated as a constant-dimension subspace packing problem. They introduce lifted Gabidulin codes to explicitly construct sparse parity-check matrices tailored for both affine and projective geometries—preserving underlying algebraic structures while effectively eliminating short cycles. The approach successfully yields sparse matrices of length up to 1024. Simulation results demonstrate a coding gain of approximately 0.5 dB over 5G LDPC codes at a block error rate of $10^{-7}$, with no evident error floor observed.

0 citationsRead paper

MetaSR: Content-Adaptive Metadata Orchestration for Generative Super-Resolution

Apr 28, 2026

This work addresses the challenge that real-world image and video content, along with degradation types, are highly complex and variable, rendering fixed metadata-guided strategies inadequate in balancing content adaptability and bandwidth constraints. To this end, we propose MetaSR, a novel framework that introduces, for the first time, a content-driven metadata orchestration mechanism. Built upon a Diffusion Transformer, MetaSR dynamically selects and injects task-relevant metadata to enable content-adaptive generative super-resolution. The method integrates a VAE and a Transformer backbone to handle heterogeneous metadata, and combines single-step diffusion distillation with rate-distortion optimization (RDO) to significantly enhance performance under limited transmission budgets. Experiments demonstrate that MetaSR achieves up to a 1.0 dB PSNR gain over baselines across diverse degradation conditions and reduces transmission bitrate by up to 50% at equivalent visual quality.

0 citationsRead paper

Harmony-Aware Music-driven Motion Synthesis with Perceptual Constraint on UGC Datasets

Jun 08, 2025

To address audio-visual desynchronization in user-generated content (UGC) dance videos—caused by misalignment between musical rhythm and human motion—the paper proposes a harmony-aware generative adversarial framework for synthesizing 3D dance motions with high rhythmic fidelity. Methodologically, it introduces a novel saliency-weighted beat evaluation strategy inspired by human visual attention, integrating cross-modal beat detection, interval-driven temporal alignment, saliency-guided beat weighting, and a unified encoder-decoder architecture enhanced with a depth refinement network. It further employs weakly supervised adversarial training stratified by beat type. Crucially, interpretable harmony modeling is embedded directly into the generation process—a first in this domain. Evaluated on limited UGC data, the method achieves statistically significant improvements over state-of-the-art approaches in both Beat Consistency Score and subjective human evaluation, yielding natural, rhythmically precise, and highly audio-visually coherent motion sequences.

0 citationsRead paper

Using In-Context Learning for Automatic Defect Labelling of Display Manufacturing Data

Jun 05, 2025

To address the high cost and low efficiency of manual defect annotation in display panel manufacturing, this paper proposes an AI-assisted automatic annotation system. Methodologically, we are the first to adapt SegGPT to industrial defect detection, introducing a domain-adaptive two-stage training paradigm—domain-specific pretraining followed by defect-aware fine-tuning—and incorporating a lightweight scribble-based annotation mechanism with a scribble-to-mask supervision strategy. Our key contributions are: (1) effective adaptation to industrial small-sample, multi-model production line data; and (2) substantial reduction in annotation dependency. Experiments on multi-model production line datasets demonstrate an average IoU improvement of 0.22, a 14% increase in recall, and an automatic annotation coverage rate of 60%. Critically, the model trained with our method achieves performance comparable to that of models trained exclusively on fully manual annotations.

0 citationsRead paper

Face Consistency Benchmark for GenAI Video

May 16, 2025

Poor cross-frame facial consistency of characters remains a critical bottleneck in AI-generated videos. This paper introduces the Facial Consistency Benchmark (FCB), the first standardized evaluation benchmark specifically designed for generative video. FCB formally defines and quantifies temporal consistency across four key facial attributes—identity, pose, expression, and illumination—and establishes a unified assessment framework integrating perceptual similarity with geometric stability. We conduct large-scale evaluations across state-of-the-art text- and image-to-video models, systematically revealing their degradation in long-term temporal consistency. FCB provides a reproducible, comparable, and diagnostic platform for model analysis and algorithmic improvement, thereby addressing a fundamental gap in the evaluation of facial consistency in generative video.

0 citationsRead paper
Recent publications

Latest Papers

Lifted Gabidulin Construction for LDPC Representations of Finite Geometry Codes

Jun 09, 2026

This work addresses the poor iterative decoding performance of low-density parity-check (LDPC) codes derived from finite geometries, which stems from the dense structure and abundance of short cycles in their conventional parity-check matrices. To overcome this limitation, the authors propose a sparsification method based on pencil selection, formulated as a constant-dimension subspace packing problem. They introduce lifted Gabidulin codes to explicitly construct sparse parity-check matrices tailored for both affine and projective geometries—preserving underlying algebraic structures while effectively eliminating short cycles. The approach successfully yields sparse matrices of length up to 1024. Simulation results demonstrate a coding gain of approximately 0.5 dB over 5G LDPC codes at a block error rate of $10^{-7}$, with no evident error floor observed.

0 citationsRead paper

MetaSR: Content-Adaptive Metadata Orchestration for Generative Super-Resolution

Apr 28, 2026

This work addresses the challenge that real-world image and video content, along with degradation types, are highly complex and variable, rendering fixed metadata-guided strategies inadequate in balancing content adaptability and bandwidth constraints. To this end, we propose MetaSR, a novel framework that introduces, for the first time, a content-driven metadata orchestration mechanism. Built upon a Diffusion Transformer, MetaSR dynamically selects and injects task-relevant metadata to enable content-adaptive generative super-resolution. The method integrates a VAE and a Transformer backbone to handle heterogeneous metadata, and combines single-step diffusion distillation with rate-distortion optimization (RDO) to significantly enhance performance under limited transmission budgets. Experiments demonstrate that MetaSR achieves up to a 1.0 dB PSNR gain over baselines across diverse degradation conditions and reduces transmission bitrate by up to 50% at equivalent visual quality.

0 citationsRead paper

Harmony-Aware Music-driven Motion Synthesis with Perceptual Constraint on UGC Datasets

Jun 08, 2025

To address audio-visual desynchronization in user-generated content (UGC) dance videos—caused by misalignment between musical rhythm and human motion—the paper proposes a harmony-aware generative adversarial framework for synthesizing 3D dance motions with high rhythmic fidelity. Methodologically, it introduces a novel saliency-weighted beat evaluation strategy inspired by human visual attention, integrating cross-modal beat detection, interval-driven temporal alignment, saliency-guided beat weighting, and a unified encoder-decoder architecture enhanced with a depth refinement network. It further employs weakly supervised adversarial training stratified by beat type. Crucially, interpretable harmony modeling is embedded directly into the generation process—a first in this domain. Evaluated on limited UGC data, the method achieves statistically significant improvements over state-of-the-art approaches in both Beat Consistency Score and subjective human evaluation, yielding natural, rhythmically precise, and highly audio-visually coherent motion sequences.

0 citationsRead paper

Using In-Context Learning for Automatic Defect Labelling of Display Manufacturing Data

Jun 05, 2025

To address the high cost and low efficiency of manual defect annotation in display panel manufacturing, this paper proposes an AI-assisted automatic annotation system. Methodologically, we are the first to adapt SegGPT to industrial defect detection, introducing a domain-adaptive two-stage training paradigm—domain-specific pretraining followed by defect-aware fine-tuning—and incorporating a lightweight scribble-based annotation mechanism with a scribble-to-mask supervision strategy. Our key contributions are: (1) effective adaptation to industrial small-sample, multi-model production line data; and (2) substantial reduction in annotation dependency. Experiments on multi-model production line datasets demonstrate an average IoU improvement of 0.22, a 14% increase in recall, and an automatic annotation coverage rate of 60%. Critically, the model trained with our method achieves performance comparable to that of models trained exclusively on fully manual annotations.

0 citationsRead paper

Face Consistency Benchmark for GenAI Video

May 16, 2025

Poor cross-frame facial consistency of characters remains a critical bottleneck in AI-generated videos. This paper introduces the Facial Consistency Benchmark (FCB), the first standardized evaluation benchmark specifically designed for generative video. FCB formally defines and quantifies temporal consistency across four key facial attributes—identity, pose, expression, and illumination—and establishes a unified assessment framework integrating perceptual similarity with geometric stability. We conduct large-scale evaluations across state-of-the-art text- and image-to-video models, systematically revealing their degradation in long-term temporal consistency. FCB provides a reproducible, comparable, and diagnostic platform for model analysis and algorithmic improvement, thereby addressing a fundamental gap in the evaluation of facial consistency in generative video.

0 citationsRead paper