Macaron-V1: 자기 개선과 Mixture-of-LoRA를 통한 오픈 지속 학습을 향하여
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
August 10, 2026
저자: Mind Lab, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yuxin Lu, Maeve Luo, Pony Ma, Verity Niu, Cole Qiao, Guian Qiu, Vince Qu, Sentry, Niko Song, Vincent Wang, Bo Wu, Rio Yang, Evelyn Ye, Fiona Ye, Ina Ye, Regis Ye, Josh Ying, Atlas Zeng, Danney Zeng, Salmon Zhan, Anya Zhang, Di Zhang, Mia Zhang, Sueky Zhang, Wei Zhao, Ada Zhou, Adrian Zhou, Yuhua Zhou, Juno Zhu, Murphy Zhuang
cs.AI
초록
Macaron-V1은 실제 환경에서의 경험을 통해 학습하고 배포 후에도 지속적으로 학습하는, 경험적 지능을 위한 오픈 에이전트-모델 패밀리이다. 이는 두 가지 시스템 목표를 중심으로 구성된다. 적응(Adaptation)은 버전 관리형 모델-하네스 쌍의 재귀적 개선을 통해 구현되며, 한 구성(configuration)에서 얻은 경험은 외부 계약(external contract) 하에서 평가되어 후속 구성을 구축하는 데 사용된다. 협업(Collaboration)은 기본 모델을 고정하고 전문 LoRA 어댑터를 구성하며 사용자 턴(turn)마다 하나의 LoRA를 선택하는 LoRA 혼합(MoL: Mixture-of-LoRA) 아키텍처를 통해 구현된다. 주력 모델인 Macaron-V1-Venti는 744B GLM-5.2 기본 모델에 채팅, 에이전트, 코딩, GenUI용 네 개의 LoRA를 결합하며, Qwen3.6 기반 Macaron-V1-Tall(50B)은 로컬 배포를 위해 동일한 설계를 사용한다. 본 보고서는 Macaron-V1을 아키텍처, 알고리즘, 인프라를 포괄하는 공동 설계 시스템으로 제시한다. MoL 아키텍처는 확장 가능한 LoRA 전문가를 통해 지속 학습을 지원한다. 알고리즘은 모델-하네스 공동 설계(Model-Harness Co-design)와 재귀적 자기 개선 루프를 결합하며, 여기에는 UI4A 컴포넌트 네이티브 GenUI 하네스, 상태 유지형 액션 서브스트레이트, 버전 관리형 HCP 계약, 에이전트형 RL 프레임워크인 MindForge가 포함된다. 지원 인프라로는 사후 학습 플랫폼 MinT, 장문맥 RL 방법 LongStraw, 희소 MoE 및 DSA 기본 모델을 위한 안정성 기법이 포함된다. 우리는 Macaron-V1을 Personal Intelligence, GenUI 및 일반 능력 벤치마크에서 최첨단 기준 모델들과 비교 평가한다. 실험 결과는 현재 시스템을 검증하며, 지속 학습과 집단 지능으로부터의 누적적 이득(compounding gains)은 여전히 미해결 과제로 남아 있다.
English
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.