Institution profile

DEEPROUTE.AI

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

In-Context Learning can Perform Continual Learning Like Humans

Sep 26, 2025

This work investigates whether large language models (LLMs) can achieve human-like continual learning via in-context learning (ICL)—specifically, retaining prior knowledge over extended multi-task sequences while accumulating new knowledge across tasks. To this end, we propose *Contextual Continual Learning* (CCL), a framework integrating task scheduling, prompt reordering, and distributed practice mechanisms, grounded in human memory similarity metrics and computational modeling of the spacing effect to mitigate catastrophic forgetting. Experiments on a Markov-chain-based multi-task benchmark demonstrate that linear-attention models (e.g., Mamba, RWKV) exhibit memory dynamics closely aligned with human behavioral patterns, including a distinct “spacing-effect sweet spot.” Crucially, CCL achieves an effective stability–plasticity trade-off without parameter updates. This is the first empirical evidence confirming that ICL—when augmented with cognitively inspired mechanisms—can support human-like continual learning.

0 citationsRead paper

NetRoller: Interfacing General and Specialized Models for End-to-End Autonomous Driving

Jun 17, 2025

To address the asynchronous collaboration challenge between large language models (LLMs) and specialized driving models (SMs)—arising from modality, temporal, and semantic discrepancies—this paper proposes NetRoller, a three-stage adapter. Its core contributions are: (1) early-stopping–based LLM semantic extraction for lightweight, inference-efficient semantic capture; (2) a hierarchical embedding scheme—comprising learnable, null, and positional embeddings—to enhance robust cross-modal alignment; and (3) Query Shift and Feature Shift mechanisms that enable low-overhead enhancement without altering the SM’s native execution frequency. Evaluated on nuScenes, NetRoller significantly improves planning performance in human-likeness (+12.7%) and safety (collision rate reduced by 28.4%), while also boosting detection and mapping accuracy by 4.3% and 5.1%, respectively. These results validate a novel paradigm for efficient, synergistic integration of general-purpose and task-specific models.

0 citationsRead paper

End-to-End HOI Reconstruction Transformer with Graph-based Encoding

Mar 08, 2025

Existing 3D HOI reconstruction methods suffer from an inherent trade-off between global structural modeling and fine-grained contact detail recovery. To address this, we propose an implicit interaction modeling paradigm that eliminates explicit interaction representation. Our method introduces a graph-residual block embedded within a Transformer architecture to jointly encode topological relationships between human and object vertices, enabling unified optimization of both global geometry and local contact regions. It integrates self-attention mechanisms, graph neural network–based encoding, differentiable mesh representations, and an end-to-end joint optimization framework. Evaluated on BEHAVE and InterCap, our approach achieves state-of-the-art performance: on InterCap, it reduces human and object mesh reconstruction errors by 8.9% and 8.6%, respectively. To the best of our knowledge, this is the first method to achieve high-fidelity, end-to-end joint mesh reconstruction for human–object interaction.

0 citationsRead paper

Car-GS: Addressing Reflective and Transparent Surface Challenges in 3D Car Reconstruction

Jan 19, 2025

To address geometric distortions in automotive 3D reconstruction caused by highly reflective (paint) and transparent (windshield) surfaces, this paper proposes the first end-to-end differentiable Gaussian splatting framework tailored for automotive scenes. Our method introduces three key innovations: (1) view-dependent Gaussian primitives to explicitly model specular reflection; (2) a learnable, geometry-rendering-decoupled opacity parameter to disentangle surface transparency from geometry estimation; and (3) a quality-aware normal supervision module that incorporates normal priors from pretrained large vision models, effectively mitigating reconstruction errors on glass under orthographic views. Evaluated on a real-world automotive dataset, our approach achieves significant improvements in normal and depth accuracy, consistently outperforming state-of-the-art methods. The resulting high-fidelity geometric representations advance applications in autonomous driving, AR, and VR.

0 citationsRead paper
Recent publications

Latest Papers

In-Context Learning can Perform Continual Learning Like Humans

Sep 26, 2025

This work investigates whether large language models (LLMs) can achieve human-like continual learning via in-context learning (ICL)—specifically, retaining prior knowledge over extended multi-task sequences while accumulating new knowledge across tasks. To this end, we propose *Contextual Continual Learning* (CCL), a framework integrating task scheduling, prompt reordering, and distributed practice mechanisms, grounded in human memory similarity metrics and computational modeling of the spacing effect to mitigate catastrophic forgetting. Experiments on a Markov-chain-based multi-task benchmark demonstrate that linear-attention models (e.g., Mamba, RWKV) exhibit memory dynamics closely aligned with human behavioral patterns, including a distinct “spacing-effect sweet spot.” Crucially, CCL achieves an effective stability–plasticity trade-off without parameter updates. This is the first empirical evidence confirming that ICL—when augmented with cognitively inspired mechanisms—can support human-like continual learning.

0 citationsRead paper

NetRoller: Interfacing General and Specialized Models for End-to-End Autonomous Driving

Jun 17, 2025

To address the asynchronous collaboration challenge between large language models (LLMs) and specialized driving models (SMs)—arising from modality, temporal, and semantic discrepancies—this paper proposes NetRoller, a three-stage adapter. Its core contributions are: (1) early-stopping–based LLM semantic extraction for lightweight, inference-efficient semantic capture; (2) a hierarchical embedding scheme—comprising learnable, null, and positional embeddings—to enhance robust cross-modal alignment; and (3) Query Shift and Feature Shift mechanisms that enable low-overhead enhancement without altering the SM’s native execution frequency. Evaluated on nuScenes, NetRoller significantly improves planning performance in human-likeness (+12.7%) and safety (collision rate reduced by 28.4%), while also boosting detection and mapping accuracy by 4.3% and 5.1%, respectively. These results validate a novel paradigm for efficient, synergistic integration of general-purpose and task-specific models.

0 citationsRead paper

End-to-End HOI Reconstruction Transformer with Graph-based Encoding

Mar 08, 2025

Existing 3D HOI reconstruction methods suffer from an inherent trade-off between global structural modeling and fine-grained contact detail recovery. To address this, we propose an implicit interaction modeling paradigm that eliminates explicit interaction representation. Our method introduces a graph-residual block embedded within a Transformer architecture to jointly encode topological relationships between human and object vertices, enabling unified optimization of both global geometry and local contact regions. It integrates self-attention mechanisms, graph neural network–based encoding, differentiable mesh representations, and an end-to-end joint optimization framework. Evaluated on BEHAVE and InterCap, our approach achieves state-of-the-art performance: on InterCap, it reduces human and object mesh reconstruction errors by 8.9% and 8.6%, respectively. To the best of our knowledge, this is the first method to achieve high-fidelity, end-to-end joint mesh reconstruction for human–object interaction.

0 citationsRead paper

Car-GS: Addressing Reflective and Transparent Surface Challenges in 3D Car Reconstruction

Jan 19, 2025

To address geometric distortions in automotive 3D reconstruction caused by highly reflective (paint) and transparent (windshield) surfaces, this paper proposes the first end-to-end differentiable Gaussian splatting framework tailored for automotive scenes. Our method introduces three key innovations: (1) view-dependent Gaussian primitives to explicitly model specular reflection; (2) a learnable, geometry-rendering-decoupled opacity parameter to disentangle surface transparency from geometry estimation; and (3) a quality-aware normal supervision module that incorporates normal priors from pretrained large vision models, effectively mitigating reconstruction errors on glass under orthographic views. Evaluated on a real-world automotive dataset, our approach achieves significant improvements in normal and depth accuracy, consistently outperforming state-of-the-art methods. The resulting high-fidelity geometric representations advance applications in autonomous driving, AR, and VR.

0 citationsRead paper