Institution profile

Mind Lab

Research institution
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Aug 10, 2026

This work addresses the challenges of continual learning in real-world deployment settings by proposing a system framework that supports open-ended self-improvement and multitask collaboration. The core innovations include a Mixture-of-LoRA (MoL) architecture that dynamically composes LoRA experts for efficient adaptation while keeping the base model frozen, a recursive co-design of model and toolchain components, and a scalable learning mechanism integrating versioned contracts, MindForge agent-based reinforcement learning, and LongStraw long-context RL. Implemented on the MinT post-training platform, the system achieves state-of-the-art performance on benchmarks spanning Personal Intelligence, GenUI, and general capabilities, demonstrating its effectiveness and scalability in open-ended continual learning scenarios.

0 citationsRead paper

CoBind: Stage-Aware Compositional Binding for Training-Free Text-to-Image Generation

Jul 14, 2026

This work addresses the challenges diffusion models face in generating images from complex text prompts involving multiple entities, attributes, and spatial relationships—often resulting in missing objects, misaligned attributes, or incorrect layouts. To tackle this, the authors propose a training-free, stage-aware compositional binding framework that parses the input prompt into a compositional graph, establishing a global layout and binding attributes during early denoising stages while progressively relaxing constraints in later stages to preserve fine details. The method innovatively incorporates dynamic guidance strength adjustment and a contrastive optimization mechanism to enable precise cross-entity attribute binding. Evaluated on benchmarks such as T2I-CompBench++ and GenEval, the approach significantly improves attribute binding accuracy, spatial relation fidelity, and complex scene generation capability, all while maintaining high visual quality.

0 citationsRead paper

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

Jun 01, 2026

This work addresses the challenge of efficiently constructing and managing massive numbers of persistent personalized models atop trillion-parameter foundation models. It proposes leveraging parameter-efficient fine-tuning (PEFT) as a lightweight and reliable personalization substrate, combining a shared large model with small, trainable adapters to encode user preferences, skills, and memory. The authors introduce MinT, an infrastructure that integrates adapter identity management, version control, provenance tracking, evaluation, and serving mechanisms, and define three scaling dimensions: Scale Up, Scale Down, and Scale Out. Experimental results demonstrate that, even under strong shared priors, compact adapters can stably capture personalized behaviors, offering a viable pathway toward large-scale deployment of millions of persistent personal models.

0 citationsRead paper

Macaron-A2UI: A Model for Generative UI in Personal Agents

May 23, 2026

Static plain-text chat interfaces struggle to support personal agents in handling complex tasks due to their lack of dynamic, context-aware interaction capabilities. This work proposes the first end-to-end generative user interface (Generative UI) model that simultaneously generates natural language responses and lightweight executable UI actions in real time—without relying on explicit structural prompts—to facilitate interactive processes such as information gathering, preference refinement, and multi-objective coordination. To advance research in this direction, we introduce a large-scale Generative UI corpus and the A2UI-Bench evaluation benchmark. Leveraging efficient LoRA-based fine-tuning and reward-driven reinforcement learning, we train large language models ranging from 30B to 754B parameters. The best-performing model achieves a score of 75.6 on A2UI-Bench, substantially outperforming current state-of-the-art baselines. All models, benchmarks, and evaluation protocols are publicly released.

0 citationsRead paper

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

May 13, 2026

This work addresses the substantial storage and serving overhead incurred by fine-tuning large language models with massive numbers of LoRA adapters by introducing the MindLab Toolkit (MinT), a hosting framework that shares a common base model while transmitting only lightweight LoRA adapters and managing their full lifecycle uniformly. Key innovations include support for catalogs of up to one million LoRA strategies, adapter compression to less than 1% of the base model size, decoupling of persistent storage from computational address spaces, and integration of tensor parallelism, GRPO-based concurrent multi-strategy training, batched MoE-LoRA loading, and cold-start-aware scheduling. Experiments demonstrate an 18.3× inference speedup on a 4B dense model and a 2.85× speedup on a 30B MoE model; a single engine can traverse 100,000 strategies, clusters support over a thousand concurrent requests, and MoE loading efficiency improves by 8.5–8.7×.

0 citationsRead paper
Recent publications

Latest Papers

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Aug 10, 2026

This work addresses the challenges of continual learning in real-world deployment settings by proposing a system framework that supports open-ended self-improvement and multitask collaboration. The core innovations include a Mixture-of-LoRA (MoL) architecture that dynamically composes LoRA experts for efficient adaptation while keeping the base model frozen, a recursive co-design of model and toolchain components, and a scalable learning mechanism integrating versioned contracts, MindForge agent-based reinforcement learning, and LongStraw long-context RL. Implemented on the MinT post-training platform, the system achieves state-of-the-art performance on benchmarks spanning Personal Intelligence, GenUI, and general capabilities, demonstrating its effectiveness and scalability in open-ended continual learning scenarios.

0 citationsRead paper

CoBind: Stage-Aware Compositional Binding for Training-Free Text-to-Image Generation

Jul 14, 2026

This work addresses the challenges diffusion models face in generating images from complex text prompts involving multiple entities, attributes, and spatial relationships—often resulting in missing objects, misaligned attributes, or incorrect layouts. To tackle this, the authors propose a training-free, stage-aware compositional binding framework that parses the input prompt into a compositional graph, establishing a global layout and binding attributes during early denoising stages while progressively relaxing constraints in later stages to preserve fine details. The method innovatively incorporates dynamic guidance strength adjustment and a contrastive optimization mechanism to enable precise cross-entity attribute binding. Evaluated on benchmarks such as T2I-CompBench++ and GenEval, the approach significantly improves attribute binding accuracy, spatial relation fidelity, and complex scene generation capability, all while maintaining high visual quality.

0 citationsRead paper

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

Jun 01, 2026

This work addresses the challenge of efficiently constructing and managing massive numbers of persistent personalized models atop trillion-parameter foundation models. It proposes leveraging parameter-efficient fine-tuning (PEFT) as a lightweight and reliable personalization substrate, combining a shared large model with small, trainable adapters to encode user preferences, skills, and memory. The authors introduce MinT, an infrastructure that integrates adapter identity management, version control, provenance tracking, evaluation, and serving mechanisms, and define three scaling dimensions: Scale Up, Scale Down, and Scale Out. Experimental results demonstrate that, even under strong shared priors, compact adapters can stably capture personalized behaviors, offering a viable pathway toward large-scale deployment of millions of persistent personal models.

0 citationsRead paper

Macaron-A2UI: A Model for Generative UI in Personal Agents

May 23, 2026

Static plain-text chat interfaces struggle to support personal agents in handling complex tasks due to their lack of dynamic, context-aware interaction capabilities. This work proposes the first end-to-end generative user interface (Generative UI) model that simultaneously generates natural language responses and lightweight executable UI actions in real time—without relying on explicit structural prompts—to facilitate interactive processes such as information gathering, preference refinement, and multi-objective coordination. To advance research in this direction, we introduce a large-scale Generative UI corpus and the A2UI-Bench evaluation benchmark. Leveraging efficient LoRA-based fine-tuning and reward-driven reinforcement learning, we train large language models ranging from 30B to 754B parameters. The best-performing model achieves a score of 75.6 on A2UI-Bench, substantially outperforming current state-of-the-art baselines. All models, benchmarks, and evaluation protocols are publicly released.

0 citationsRead paper

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

May 13, 2026

This work addresses the substantial storage and serving overhead incurred by fine-tuning large language models with massive numbers of LoRA adapters by introducing the MindLab Toolkit (MinT), a hosting framework that shares a common base model while transmitting only lightweight LoRA adapters and managing their full lifecycle uniformly. Key innovations include support for catalogs of up to one million LoRA strategies, adapter compression to less than 1% of the base model size, decoupling of persistent storage from computational address spaces, and integration of tensor parallelism, GRPO-based concurrent multi-strategy training, batched MoE-LoRA loading, and cold-start-aware scheduling. Experiments demonstrate an 18.3× inference speedup on a 4B dense model and a 2.85× speedup on a 30B MoE model; a single engine can traverse 100,000 strategies, clusters support over a thousand concurrent requests, and MoE loading efficiency improves by 8.5–8.7×.

0 citationsRead paper