Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yuxin Lu, Maeve Luo, Pony Ma, Verity Niu, Cole Qiao, Guian Qiu, Vince Qu, Sentry, Niko Song, Vincent Wang, Bo Wu, Rio Yang, Evelyn Ye, Fiona Ye, Ina Ye, Regis Ye, Josh Ying, Atlas Zeng, Danney Zeng, Salmon Zhan, Anya Zhang, Di Zhang, Mia Zhang, Sueky Zhang, Wei Zhao, Ada Zhou, Adrian Zhou, Yuhua Zhou, Juno Zhu, Murphy Zhuang

arXiv:2608.09819 · 2026-08-11 공개 · arXiv · PDF

continual-learning sparse-moe glm-5-2 agent-model mixture-of-lora longstraw min-t dsa

Abstract

Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.

한국어 요약

한 줄 요약

Macaron-V1은 경험 학습과 협업을 위한 개방형 지속 학습 시스템으로, Mixture-of-LoRA와 자기 개선 루프를 결합한 모델이다.

핵심 기여도

핵심 아이디어

Macaron-V1은 기존 모델이 배포 후에도 경험을 통해 지속적으로 개선되도록 설계된 시스템이다. 이는 기반 모델을 고정하고, 다양한 LoRA 어댑터를 통해 특정 작업에 맞게 유연하게 조합하는 Mixture-of-LoRA (MoL) 아키텍처를 통해 실현된다. 또한, Model-Harness Co-design과 recursive self-improvement 루프를 통해 모델과 헤이서스가 함께 개선되도록 유도한다. UI4A와 GenUI 헤이서스, MindForge RL 프레임워크 등은 사용자 경험과 행동 기반 학습을 구현하는 핵심 구성 요소이다.

기술적 접근법

주요 결과

의의 및 한계

Macaron-V1은 경험 기반 지능과 지속 학습을 위한 체계적인 시스템 설계를 제시하며, 모델-헤이서스 공동 설계와 자기 개선 루프를 통해 유연성과 확장성을 높였다. 그러나 지속 학습의 복합적 이득과 집단 지능의 효과는 여전히 개방된 질문으로 남아 있다. 또한, MoL 아키텍처는 LoRA 어댑터의 관리와 선택에 대한 추가 연구가 필요할 수 있다.

실용적 활용

Macaron-V1은 개인화된 인공지능 서비스, UI 기반 작업 자동화, 그리고 로컬 환경에서의 대규모 모델 배포에 적용 가능하다. 특히, 지속 학습 기반의 에이전트 시스템 개발에 유용할 것으로 기대된다.