Institution profile

Improbable

Industry researcheurope · gb
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Self-Distillation Enables Continual Learning

Jan 27, 2026

This work addresses the challenges of catastrophic forgetting and the lack of effective online policy learning methods in continual learning for foundation models. To overcome these limitations, the authors propose Self-Distillation Fine-Tuning (SDFT), a novel approach that enables online policy self-distillation using only expert demonstrations. SDFT leverages in-context learning to generate training data by treating the model’s own outputs—conditioned on provided demonstrations—as teacher signals, thereby eliminating the need for explicit reward functions. This mechanism simultaneously preserves previously acquired knowledge while acquiring new skills. Experimental results demonstrate that SDFT substantially outperforms conventional supervised fine-tuning, achieving superior performance on new tasks, effectively mitigating catastrophic forgetting, and enabling the sequential accumulation of multiple skills without performance degradation.

5 citationsRead paper

Training Language Models via Neural Cellular Automata

Mar 09, 2026

This work addresses key limitations in natural language pretraining—such as data scarcity, human biases, and the entanglement of knowledge with reasoning—by proposing a novel approach that leverages neural cellular automata (NCA) to generate synthetic, non-linguistic data exhibiting language-like statistical properties. The method introduces NCA-generated sequences for pre-pretraining large language models before fine-tuning on real textual data. This is the first application of NCAs to language model pretraining, offering scalable, controllable synthetic data that can be tailored in complexity to specific target domains. Using only 164 million NCA tokens for pre-pretraining, the approach achieves up to a 6% improvement in downstream language modeling performance, accelerates convergence by 1.6×, and significantly outperforms baselines on reasoning benchmarks including GSM8K, HumanEval, and BigBench-Lite.

0 citationsRead paper

Going Beyond Heuristics by Imposing Policy Improvement as a Constraint

Jul 07, 2025

In reinforcement learning, hand-crafted heuristic rewards are commonly used to inject domain knowledge; however, their misalignment with the task reward often causes performance degradation, and existing “policy-invariance”-based approaches lack theoretical guarantees of improvement. To address this, we propose HEPO—a framework that abandons the policy-invariance assumption and instead enforces explicit policy improvement constraints, dynamically balancing task objectives and prior knowledge using heuristic rewards as auxiliary signals. HEPO effectively mitigates reward hacking, accommodates coarse, non-expert-designed heuristics, and substantially reduces reward engineering overhead. Empirical evaluation on standard benchmarks demonstrates that HEPO outperforms state-of-the-art methods, maintains robustness and efficiency even under low-quality heuristics, and enhances the practicality and scalability of reward design in RL.

0 citationsRead paper

Self-Adapting Language Models

Jun 12, 2025

Current large language models (LLMs) lack the capability to autonomously update their parameters to adapt to new tasks or incorporate novel knowledge. Method: This paper introduces the first purely generative, end-to-end self-adaptation training paradigm, wherein the model autonomously generates editing instructions and synthesizes fine-tuning data to drive persistent parameter updates—without external modules or human intervention. The approach integrates supervised fine-tuning (SFT), reinforcement learning with reward derived from post-update performance, tool invocation, and self-supervised data augmentation. Contribution/Results: Our method achieves significant improvements in knowledge injection and few-shot generalization, demonstrating robust adaptation efficacy. Crucially, it provides the first empirical validation that LLMs possess an intrinsic capacity for self-directed, continual evolution—establishing a foundational step toward autonomous model adaptation.

0 citationsRead paper
Recent publications

Latest Papers

Training Language Models via Neural Cellular Automata

Mar 09, 2026

This work addresses key limitations in natural language pretraining—such as data scarcity, human biases, and the entanglement of knowledge with reasoning—by proposing a novel approach that leverages neural cellular automata (NCA) to generate synthetic, non-linguistic data exhibiting language-like statistical properties. The method introduces NCA-generated sequences for pre-pretraining large language models before fine-tuning on real textual data. This is the first application of NCAs to language model pretraining, offering scalable, controllable synthetic data that can be tailored in complexity to specific target domains. Using only 164 million NCA tokens for pre-pretraining, the approach achieves up to a 6% improvement in downstream language modeling performance, accelerates convergence by 1.6×, and significantly outperforms baselines on reasoning benchmarks including GSM8K, HumanEval, and BigBench-Lite.

0 citationsRead paper

Self-Distillation Enables Continual Learning

Jan 27, 2026

This work addresses the challenges of catastrophic forgetting and the lack of effective online policy learning methods in continual learning for foundation models. To overcome these limitations, the authors propose Self-Distillation Fine-Tuning (SDFT), a novel approach that enables online policy self-distillation using only expert demonstrations. SDFT leverages in-context learning to generate training data by treating the model’s own outputs—conditioned on provided demonstrations—as teacher signals, thereby eliminating the need for explicit reward functions. This mechanism simultaneously preserves previously acquired knowledge while acquiring new skills. Experimental results demonstrate that SDFT substantially outperforms conventional supervised fine-tuning, achieving superior performance on new tasks, effectively mitigating catastrophic forgetting, and enabling the sequential accumulation of multiple skills without performance degradation.

5 citationsRead paper

Going Beyond Heuristics by Imposing Policy Improvement as a Constraint

Jul 07, 2025

In reinforcement learning, hand-crafted heuristic rewards are commonly used to inject domain knowledge; however, their misalignment with the task reward often causes performance degradation, and existing “policy-invariance”-based approaches lack theoretical guarantees of improvement. To address this, we propose HEPO—a framework that abandons the policy-invariance assumption and instead enforces explicit policy improvement constraints, dynamically balancing task objectives and prior knowledge using heuristic rewards as auxiliary signals. HEPO effectively mitigates reward hacking, accommodates coarse, non-expert-designed heuristics, and substantially reduces reward engineering overhead. Empirical evaluation on standard benchmarks demonstrates that HEPO outperforms state-of-the-art methods, maintains robustness and efficiency even under low-quality heuristics, and enhances the practicality and scalability of reward design in RL.

0 citationsRead paper

Self-Adapting Language Models

Jun 12, 2025

Current large language models (LLMs) lack the capability to autonomously update their parameters to adapt to new tasks or incorporate novel knowledge. Method: This paper introduces the first purely generative, end-to-end self-adaptation training paradigm, wherein the model autonomously generates editing instructions and synthesizes fine-tuning data to drive persistent parameter updates—without external modules or human intervention. The approach integrates supervised fine-tuning (SFT), reinforcement learning with reward derived from post-update performance, tool invocation, and self-supervised data augmentation. Contribution/Results: Our method achieves significant improvements in knowledge injection and few-shot generalization, demonstrating robust adaptation efficacy. Crucially, it provides the first empirical validation that LLMs possess an intrinsic capacity for self-directed, continual evolution—establishing a foundational step toward autonomous model adaptation.

0 citationsRead paper