Beyond Embedding Transfer: Component Roles in Grokking Transfer and Stability
研究通过转移内部注意力/MLP权重及读出组件来加速算法任务的泛化,并探讨了继续优化时的稳定性问题。
研究通过转移内部注意力/MLP权重及读出组件来加速算法任务的泛化,并探讨了继续优化时的稳定性问题。
Molecular force prediction is often hindered by a mismatch between predefined spatial scales and the optimal modeling scale for a given task. This work proposes a loss-guided adaptive scale optimization framework that, for the first time, leverages loss signals to drive multi-scale selection in molecular representation learning, thereby transcending conventional fixed-scale modeling paradigms. The approach integrates hard routing, continuous interpolation, differentiable scale updating, and a dynamic scale-pool refinement mechanism to automatically search for an improved modeling resolution starting from an initial anchor scale. Evaluated on an aqueous NaCl system, the method reduces the overall force prediction mean absolute error (MAE) to 381.23 meV/Å, with a marked improvement in the near-contact region (<0.6 nm), where the MAE drops to 260.51 meV/Å.
Rare diseases suffer from prolonged diagnostic timelines and fragmented clinical evidence; moreover, general-purpose large language models exhibit limited clinical reasoning capabilities due to scarce real-world electronic health records (EHRs), outdated medical knowledge, and hallucination. To address these challenges, we propose a domain-specific clinical reasoning paradigm centered on “narrative-first, knowledge-enhanced” inference. Our approach comprises: (1) constructing a physician-validated rare-disease reasoning dataset and domain-specific corpus; (2) designing a knowledge graph–anchored retrieval mechanism and phased chain-of-thought training to integrate non-phenotypic evidence (e.g., imaging, functional tests); and (3) enhancing robustness under noisy conditions and phenotypic overlap via instruction fine-tuning, knowledge graph fusion, and structured reasoning. Evaluated on multicenter real-world EHRs and public benchmarks, our method achieves state-of-the-art performance—matching the diagnostic accuracy of senior clinicians while significantly shortening diagnostic pathways and enabling transparent, auditable clinical decision-making.
研究通过转移内部注意力/MLP权重及读出组件来加速算法任务的泛化,并探讨了继续优化时的稳定性问题。
Molecular force prediction is often hindered by a mismatch between predefined spatial scales and the optimal modeling scale for a given task. This work proposes a loss-guided adaptive scale optimization framework that, for the first time, leverages loss signals to drive multi-scale selection in molecular representation learning, thereby transcending conventional fixed-scale modeling paradigms. The approach integrates hard routing, continuous interpolation, differentiable scale updating, and a dynamic scale-pool refinement mechanism to automatically search for an improved modeling resolution starting from an initial anchor scale. Evaluated on an aqueous NaCl system, the method reduces the overall force prediction mean absolute error (MAE) to 381.23 meV/Å, with a marked improvement in the near-contact region (<0.6 nm), where the MAE drops to 260.51 meV/Å.
Rare diseases suffer from prolonged diagnostic timelines and fragmented clinical evidence; moreover, general-purpose large language models exhibit limited clinical reasoning capabilities due to scarce real-world electronic health records (EHRs), outdated medical knowledge, and hallucination. To address these challenges, we propose a domain-specific clinical reasoning paradigm centered on “narrative-first, knowledge-enhanced” inference. Our approach comprises: (1) constructing a physician-validated rare-disease reasoning dataset and domain-specific corpus; (2) designing a knowledge graph–anchored retrieval mechanism and phased chain-of-thought training to integrate non-phenotypic evidence (e.g., imaging, functional tests); and (3) enhancing robustness under noisy conditions and phenotypic overlap via instruction fine-tuning, knowledge graph fusion, and structured reasoning. Evaluated on multicenter real-world EHRs and public benchmarks, our method achieves state-of-the-art performance—matching the diagnostic accuracy of senior clinicians while significantly shortening diagnostic pathways and enabling transparent, auditable clinical decision-making.