Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of continual learning in real-world deployment settings by proposing a system framework that supports open-ended self-improvement and multitask collaboration. The core innovations include a Mixture-of-LoRA (MoL) architecture that dynamically composes LoRA experts for efficient adaptation while keeping the base model frozen, a recursive co-design of model and toolchain components, and a scalable learning mechanism integrating versioned contracts, MindForge agent-based reinforcement learning, and LongStraw long-context RL. Implemented on the MinT post-training platform, the system achieves state-of-the-art performance on benchmarks spanning Personal Intelligence, GenUI, and general capabilities, demonstrating its effectiveness and scalability in open-ended continual learning scenarios.
📝 Abstract
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.
Problem

Research questions and friction points this paper is trying to address.

open continual learning
self-improvement
Mixture-of-LoRA
experiential intelligence
agent-model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-LoRA
continual learning
Model-Harness Co-design
recursive self-improvement
experiential intelligence
🔎 Similar Papers
2023-12-01arXiv.orgCitations: 3
💼 Related Jobs
No related jobs found.
V
Vin Bo
Mind Lab
A
Asher Cai
Mind Lab
J
Jingwei Cao
Mind Lab
Song Cao
Song Cao
University of Southern California
Computer Vision
V
Vic Cao
Mind Lab
A
Amelia Chen
Mind Lab
A
Andrew Chen
Mind Lab
K
Kaijie Chen
Mind Lab
C
Cleon Cheng
Mind Lab
S
Steven Chiang
Mind Lab
Kaixuan Fan
Kaixuan Fan
MSc Student, Imperial College London
Virtual RealityHuman-Computer InteractionComputer Vision
H
Hera Feng
Mind Lab
H
Huan Feng
Mind Lab
A
Arthur Fu
Mind Lab
J
Jun Gao
Mind Lab
P
Pyke Han
Mind Lab
N
Nolan Ho
Mind Lab
O
Ori Hong
Mind Lab
H
Hailee Hou
Mind Lab
P
Piers Hua
Mind Lab
C
Charles Huang
Mind Lab
M
Miles Jiang
Mind Lab
N
Nora Jiang
Mind Lab