Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
This work addresses the challenges of continual learning in real-world deployment settings by proposing a system framework that supports open-ended self-improvement and multitask collaboration. The core innovations include a Mixture-of-LoRA (MoL) architecture that dynamically composes LoRA experts for efficient adaptation while keeping the base model frozen, a recursive co-design of model and toolchain components, and a scalable learning mechanism integrating versioned contracts, MindForge agent-based reinforcement learning, and LongStraw long-context RL. Implemented on the MinT post-training platform, the system achieves state-of-the-art performance on benchmarks spanning Personal Intelligence, GenUI, and general capabilities, demonstrating its effectiveness and scalability in open-ended continual learning scenarios.