Institution profile

China Nanhu Academy of Electronics and Information Technology

Academic institutionasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

May 11, 2026

This work addresses the lack of effective, interpretable, and low-cost guidance for determining layer-wise freezing or training strategies during continual pre-training of large language models. The authors propose LayerTracer, a framework that identifies task-executing layers, quantifies layer sensitivity, and analyzes the stability of representational evolution across layers. Their analysis reveals that deeper layers are primarily responsible for task execution and exhibit greater robustness to perturbations. Based on these insights, they introduce an interpretable and cost-efficient continual pre-training paradigm—freezing deeper layers while fine-tuning shallower ones—and extend it to hybrid model construction. The approach is architecture-agnostic and, through representation analysis and controlled experiments, consistently outperforms full-parameter fine-tuning and alternative strategies on C-Eval and CMMLU benchmarks, while demonstrating that embedding high-quality deep modules effectively preserves original knowledge.

0 citationsRead paper

LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures

Apr 22, 2026

This study addresses the limited understanding of hierarchical representation dynamics, task-specific knowledge localization, and robustness bottlenecks in diverse large language model architectures, which hinders effective hybrid architecture design and optimization. To this end, we propose LayerTracer, a framework that jointly defines and analyzes “task particles”—layers where target token probabilities exhibit significant increases—and “fragile layers”—those whose output distributions undergo maximal shifts under perturbation. By sequentially extracting hidden states, mapping them to vocabulary probability distributions, and quantifying perturbation effects via masked interventions and Jensen–Shannon divergence, LayerTracer provides an architecture-agnostic analytical tool. Experiments reveal that task particles predominantly emerge in deeper layers and that larger models exhibit stronger layer-wise robustness. The framework enables precise identification of critical layers, offering principled guidance for layer partitioning, module allocation, and gating strategies in hybrid models.

0 citationsRead paper
Recent publications

Latest Papers

Freeze Deep, Train Shallow: Interpretable Layer Allocation for Continued Pre-Training

May 11, 2026

This work addresses the lack of effective, interpretable, and low-cost guidance for determining layer-wise freezing or training strategies during continual pre-training of large language models. The authors propose LayerTracer, a framework that identifies task-executing layers, quantifies layer sensitivity, and analyzes the stability of representational evolution across layers. Their analysis reveals that deeper layers are primarily responsible for task execution and exhibit greater robustness to perturbations. Based on these insights, they introduce an interpretable and cost-efficient continual pre-training paradigm—freezing deeper layers while fine-tuning shallower ones—and extend it to hybrid model construction. The approach is architecture-agnostic and, through representation analysis and controlled experiments, consistently outperforms full-parameter fine-tuning and alternative strategies on C-Eval and CMMLU benchmarks, while demonstrating that embedding high-quality deep modules effectively preserves original knowledge.

0 citationsRead paper

LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures

Apr 22, 2026

This study addresses the limited understanding of hierarchical representation dynamics, task-specific knowledge localization, and robustness bottlenecks in diverse large language model architectures, which hinders effective hybrid architecture design and optimization. To this end, we propose LayerTracer, a framework that jointly defines and analyzes “task particles”—layers where target token probabilities exhibit significant increases—and “fragile layers”—those whose output distributions undergo maximal shifts under perturbation. By sequentially extracting hidden states, mapping them to vocabulary probability distributions, and quantifying perturbation effects via masked interventions and Jensen–Shannon divergence, LayerTracer provides an architecture-agnostic analytical tool. Experiments reveal that task particles predominantly emerge in deeper layers and that larger models exhibit stronger layer-wise robustness. The framework enables precise identification of critical layers, offering principled guidance for layer partitioning, module allocation, and gating strategies in hybrid models.

0 citationsRead paper