Institution profile

IAI MIPT LLC

Industry researcheurope · ru
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

Oct 29, 2025

This work addresses the degradation of vision-language (VL) representations during action fine-tuning of Vision-Language-Action (VLA) models, which impairs out-of-distribution (OOD) generalization. We systematically characterize the trade-off between action adaptation and visual representation collapse. To mitigate this, we propose a lightweight hidden-layer representation alignment strategy that explicitly preserves pre-trained VL knowledge via cross-task feature constraints and attention-guided regularization—without incurring additional inference overhead. Through representation probing, attention visualization, and ablation on contrastive tasks, we demonstrate that our method significantly alleviates visual representation degradation. Empirically, it improves OOD generalization across multiple robotic manipulation benchmarks. The implementation is publicly available.

0 citationsRead paper

A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control

Oct 15, 2025

Prior work has identified instability in Transformer-based online model-free reinforcement learning (RL), primarily due to high sensitivity to policy/value network architecture, parameter sharing schemes, and temporal modeling strategies. Method: This paper presents the first systematic study of Transformer design for online continuous control, proposing a stable and efficient Actor-Critic architecture featuring serialized state inputs, temporal slicing, cross-network parameter sharing, and conditional input conditioning—unified to support both vector and image observations. Contribution/Results: The proposed method significantly improves training stability and generalization across diverse online RL benchmarks. It achieves state-of-the-art performance on both fully observed (e.g., MuJoCo) and partially observed (e.g., DeepMind Control Suite with proprioceptive+visual inputs) tasks. By providing a reproducible architectural blueprint and empirically validated design principles, this work establishes a new paradigm and practical guidelines for deploying Transformers in online RL settings.

0 citationsRead paper

ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL

Oct 08, 2025

Real-world robotics demands decision-making under partial observability and long-horizon dependencies, yet existing methods—constrained by fixed attention windows or unstable memory mechanisms—struggle to model ultra-long historical dependencies. To address this, we propose a hierarchical external memory architecture comprising localized layer-wise memory modules, bidirectional cross-attention, and a convex mixture update/rewrite strategy based on Linear Recurrent Units (LRUs). This design overcomes context-length limitations while enabling efficient, stable long-term memory maintenance. Crucially, it integrates seamlessly into the Transformer framework without inflating sequence length, supporting dependency modeling over million-step trajectories. Empirically, our method achieves 100% success on T-Maze and significantly outperforms state-of-the-art baselines on POPGym and MIKASA-Robo visual manipulation tasks, demonstrating substantial improvement in history modeling for partially observable reinforcement learning.

0 citationsRead paper

Accelerating Transformers in Online RL

Sep 30, 2025

Transformer-based policies in model-free online reinforcement learning suffer from training instability, slow convergence, and heavy reliance on large replay buffers. To address these challenges, this work proposes a two-stage training framework: first, stabilizing Transformer policy initialization via behavior cloning using a foundational model (e.g., a pretrained policy network); second, switching to fully online, interactive RL fine-tuning. Crucially, the foundational model serves as a “training accelerator,” mitigating optimization difficulties inherent to Transformers in online settings. Experiments on ManiSkill (vision-based, POMDP) and MuJoCo (state-based, MDP) benchmarks demonstrate that our approach doubles training speed in visual domains, reduces replay buffer requirements to only 10–20k transitions, and significantly lowers computational overhead—while preserving generalization capability and training stability.

0 citationsRead paper
Recent publications

Latest Papers

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

Oct 29, 2025

This work addresses the degradation of vision-language (VL) representations during action fine-tuning of Vision-Language-Action (VLA) models, which impairs out-of-distribution (OOD) generalization. We systematically characterize the trade-off between action adaptation and visual representation collapse. To mitigate this, we propose a lightweight hidden-layer representation alignment strategy that explicitly preserves pre-trained VL knowledge via cross-task feature constraints and attention-guided regularization—without incurring additional inference overhead. Through representation probing, attention visualization, and ablation on contrastive tasks, we demonstrate that our method significantly alleviates visual representation degradation. Empirically, it improves OOD generalization across multiple robotic manipulation benchmarks. The implementation is publicly available.

0 citationsRead paper

A New Perspective on Transformers in Online Reinforcement Learning for Continuous Control

Oct 15, 2025

Prior work has identified instability in Transformer-based online model-free reinforcement learning (RL), primarily due to high sensitivity to policy/value network architecture, parameter sharing schemes, and temporal modeling strategies. Method: This paper presents the first systematic study of Transformer design for online continuous control, proposing a stable and efficient Actor-Critic architecture featuring serialized state inputs, temporal slicing, cross-network parameter sharing, and conditional input conditioning—unified to support both vector and image observations. Contribution/Results: The proposed method significantly improves training stability and generalization across diverse online RL benchmarks. It achieves state-of-the-art performance on both fully observed (e.g., MuJoCo) and partially observed (e.g., DeepMind Control Suite with proprioceptive+visual inputs) tasks. By providing a reproducible architectural blueprint and empirically validated design principles, this work establishes a new paradigm and practical guidelines for deploying Transformers in online RL settings.

0 citationsRead paper

ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL

Oct 08, 2025

Real-world robotics demands decision-making under partial observability and long-horizon dependencies, yet existing methods—constrained by fixed attention windows or unstable memory mechanisms—struggle to model ultra-long historical dependencies. To address this, we propose a hierarchical external memory architecture comprising localized layer-wise memory modules, bidirectional cross-attention, and a convex mixture update/rewrite strategy based on Linear Recurrent Units (LRUs). This design overcomes context-length limitations while enabling efficient, stable long-term memory maintenance. Crucially, it integrates seamlessly into the Transformer framework without inflating sequence length, supporting dependency modeling over million-step trajectories. Empirically, our method achieves 100% success on T-Maze and significantly outperforms state-of-the-art baselines on POPGym and MIKASA-Robo visual manipulation tasks, demonstrating substantial improvement in history modeling for partially observable reinforcement learning.

0 citationsRead paper

Accelerating Transformers in Online RL

Sep 30, 2025

Transformer-based policies in model-free online reinforcement learning suffer from training instability, slow convergence, and heavy reliance on large replay buffers. To address these challenges, this work proposes a two-stage training framework: first, stabilizing Transformer policy initialization via behavior cloning using a foundational model (e.g., a pretrained policy network); second, switching to fully online, interactive RL fine-tuning. Crucially, the foundational model serves as a “training accelerator,” mitigating optimization difficulties inherent to Transformers in online settings. Experiments on ManiSkill (vision-based, POMDP) and MuJoCo (state-based, MDP) benchmarks demonstrate that our approach doubles training speed in visual domains, reduces replay buffer requirements to only 10–20k transitions, and significantly lowers computational overhead—while preserving generalization capability and training stability.

0 citationsRead paper