Institution profile

TeleAI

Industry researchasia · cn
Official website
Research library120linked papers
Opportunities0open roles
Selected work

Representative Papers

Dual Attention Residuals

Jul 21, 2026

This work addresses the limited dynamic depth selection in conventional Transformer residual architectures, which stems from insufficient historical information exchange across multiple parallel streams. To overcome this, we propose a reciprocal cross-stream addressing mechanism that enables bidirectional historical retrieval within a multi-stream framework: each stream computes depth weights based on the state of its counterpart stream and applies these weights to its own historical values. Our approach uniquely integrates inter-stream interaction into historical retrieval, preserving inter-layer representational diversity through cross-stream depth selection while mitigating redundancy and functional imbalance. Key components include reciprocal cross-attention, normalized state weighting, constrained gated writing, and block-level history storage. Experiments demonstrate consistent and significant improvements over standard residual Transformers and Attention Residuals across dense models (0.1B–1B) and a 7B sparse MoE model, with ablation studies confirming that performance gains arise from cross-stream interaction rather than additional parameters or projections.

0 citationsRead paper

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation

Jul 13, 2026

This work addresses a key limitation in existing robotic manipulation approaches, which often neglect the dependence of action semantics on environmental context, resulting in noisy, redundant, and poorly structured control trajectories. To overcome this, the paper introduces EDAR—a novel framework that explicitly models the coupling between actions and their surrounding context. EDAR constructs environment-dependent action tokens by jointly embedding executable control commands with their visual outcomes in specific scenes. This representation enables the action space to capture interaction semantics rather than merely encoding command patterns. Evaluated in both simulated and real-world robotic manipulation tasks, EDAR significantly enhances downstream policy learning performance, demonstrating particularly strong gains in long-horizon manipulation scenarios.

0 citationsRead paper

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Jun 22, 2026

Existing real-time streaming methods for virtual human video generation struggle to simultaneously maintain long-term visual temporal consistency and accurately perceive user intent. This work proposes a real-time framework capable of generating videos of unlimited duration, leveraging autoregressive distillation to enhance inference efficiency. It introduces an innovative fusion of short- and long-term visual memory mechanisms with a reasoning-and-response module, complemented by a state-recurrent strategy and a cache-switching mechanism. This design enables high visual consistency while effectively aligning with complex user intentions. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across diverse scenarios, achieving—for the first time in real-time streaming generation—concurrent long-term visual coherence and responsive interactive intent alignment.

0 citationsRead paper
Recent publications

Latest Papers

Dual Attention Residuals

Jul 21, 2026

This work addresses the limited dynamic depth selection in conventional Transformer residual architectures, which stems from insufficient historical information exchange across multiple parallel streams. To overcome this, we propose a reciprocal cross-stream addressing mechanism that enables bidirectional historical retrieval within a multi-stream framework: each stream computes depth weights based on the state of its counterpart stream and applies these weights to its own historical values. Our approach uniquely integrates inter-stream interaction into historical retrieval, preserving inter-layer representational diversity through cross-stream depth selection while mitigating redundancy and functional imbalance. Key components include reciprocal cross-attention, normalized state weighting, constrained gated writing, and block-level history storage. Experiments demonstrate consistent and significant improvements over standard residual Transformers and Attention Residuals across dense models (0.1B–1B) and a 7B sparse MoE model, with ablation studies confirming that performance gains arise from cross-stream interaction rather than additional parameters or projections.

0 citationsRead paper

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation

Jul 13, 2026

This work addresses a key limitation in existing robotic manipulation approaches, which often neglect the dependence of action semantics on environmental context, resulting in noisy, redundant, and poorly structured control trajectories. To overcome this, the paper introduces EDAR—a novel framework that explicitly models the coupling between actions and their surrounding context. EDAR constructs environment-dependent action tokens by jointly embedding executable control commands with their visual outcomes in specific scenes. This representation enables the action space to capture interaction semantics rather than merely encoding command patterns. Evaluated in both simulated and real-world robotic manipulation tasks, EDAR significantly enhances downstream policy learning performance, demonstrating particularly strong gains in long-horizon manipulation scenarios.

0 citationsRead paper

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars

Jun 22, 2026

Existing real-time streaming methods for virtual human video generation struggle to simultaneously maintain long-term visual temporal consistency and accurately perceive user intent. This work proposes a real-time framework capable of generating videos of unlimited duration, leveraging autoregressive distillation to enhance inference efficiency. It introduces an innovative fusion of short- and long-term visual memory mechanisms with a reasoning-and-response module, complemented by a state-recurrent strategy and a cache-switching mechanism. This design enables high visual consistency while effectively aligning with complex user intentions. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches across diverse scenarios, achieving—for the first time in real-time streaming generation—concurrent long-term visual coherence and responsive interactive intent alignment.

0 citationsRead paper