Institution profile

China Telecom

Industry researchasia · cn
Official website
Research library285linked papers
Opportunities0open roles
Selected work

Representative Papers

Style-CCL: Content-Preserving Style Transfer via Curriculum Continual Learning

Jun 04, 2026

Existing diffusion Transformers struggle to effectively disentangle content and style in style transfer, often being dominated by semantic-level style cues, which leads to insufficient texture learning and content distortion. To address this, this work proposes Style-CCL—the first multi-stage training framework that integrates curriculum learning with continual learning. It trains a dual-branch SC-DiT model following a structured progression from “semantic to texture” and “clean to synthetic” data, while incorporating random memory replay to mitigate catastrophic forgetting. The approach leverages independent RoPE embeddings, causal masking, and reverse triplets to construct a million-scale dataset. Evaluated on style similarity, content consistency, and aesthetic quality, Style-CCL achieves state-of-the-art performance, significantly enhancing both stylistic expressiveness and content fidelity.

1 citationsRead paper

TextOp: Real-time Interactive Text-Driven Humanoid Robot Motion Generation and Control

Feb 07, 2026

This work addresses the challenge of driving general-purpose humanoid robots in real time with interactive, flexible user intent while enabling autonomous execution. To this end, the authors propose TextOp, a framework that combines a high-level autoregressive motion diffusion model to generate short-horizon full-body motion trajectories from streaming text instructions in real time, and a low-level robust tracking policy to accurately execute these motions on physical robots. TextOp is the first method to support dynamic modification of instructions during execution, enabling free-form intent expression and seamless switching among multiple behaviors. Experiments on real robots demonstrate the system’s immediate responsiveness, motion smoothness, and control precision, successfully achieving fluid transitions between complex actions such as dancing and jumping.

1 citationsRead paper

Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence

Jan 21, 2026

This work addresses the challenge of deploying visual surveillance in privacy-sensitive areas—such as restrooms and changing rooms—where ethical and regulatory constraints prohibit conventional camera-based systems. The authors propose the Single-Pixel Vision-Language Model (SP-VLM), which, for the first time, integrates single-pixel sensing with multimodal vision-language fusion to enable high-level semantic understanding of human activities from extremely low-dimensional observations. By design, SP-VLM inherently preserves privacy by effectively preventing identity reconstruction while simultaneously supporting tasks such as anomaly detection, occupancy counting, and activity recognition. Experimental results demonstrate that, below a critical sampling rate where mainstream face recognition systems fail completely, SP-VLM maintains high accuracy, thereby establishing a novel paradigm that harmonizes privacy preservation with intelligent behavioral analysis.

1 citationsRead paper

GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

Jan 14, 2026

This work proposes GUI-Eyes, a reinforcement learning–based active visual perception framework that addresses the limitations of existing GUI agents, which predominantly rely on static visual inputs and lack the ability to actively decide when and how to observe the interface. GUI-Eyes employs a two-stage policy network to autonomously invoke visual tools—such as cropping and zooming—to enable progressive perception from coarse exploration to fine-grained localization. The approach innovatively integrates a vision-language model with a tool-calling mechanism and introduces a continuous-space reward function that combines positional proximity and region overlap, effectively mitigating reward sparsity in GUI environments. Evaluated on the ScreenSpot-Pro benchmark, GUI-Eyes-3B achieves a localization accuracy of 44.8% using only 3k annotated samples, significantly outperforming current supervised and reinforcement learning baselines.

1 citationsRead paper

QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit

Jan 08, 2026arXiv.org

This work addresses the challenge of effectively disentangling content and style in content-preserving style transfer using diffusion transformers (DiTs), a task where existing methods often suffer from content distortion or stylistic inaccuracies. We propose QwenStyle V1, the first DiT-based model for this task, unlocking the potential of Qwen-Image-Edit in high-fidelity style transfer. By synthesizing high-quality in-the-wild style triplets and incorporating a curriculum continual learning framework, our model generalizes robustly to unseen styles—even when trained on a mixture of clean and noisy data—while preserving content fidelity. Extensive experiments demonstrate that QwenStyle V1 achieves state-of-the-art performance across three key metrics: style similarity, content consistency, and aesthetic quality.

1 citationsRead paper
Recent publications

Latest Papers