ChatPaper.aiChatPaper

Macaron-V1:邁向具備自我改進與LoRA混合的開放式持續學習

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

August 10, 2026
作者: Mind Lab, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yuxin Lu, Maeve Luo, Pony Ma, Verity Niu, Cole Qiao, Guian Qiu, Vince Qu, Sentry, Niko Song, Vincent Wang, Bo Wu, Rio Yang, Evelyn Ye, Fiona Ye, Ina Ye, Regis Ye, Josh Ying, Atlas Zeng, Danney Zeng, Salmon Zhan, Anya Zhang, Di Zhang, Mia Zhang, Sueky Zhang, Wei Zhao, Ada Zhou, Adrian Zhou, Yuhua Zhou, Juno Zhu, Murphy Zhuang
cs.AI

摘要

Macaron-V1 是一個面向經驗智能的開放式智能體模型系列:從真實環境中的經驗學習,並在部署後持續學習。其設計圍繞兩個系統目標展開。適應性透過對版本化模型-測試框架配對的遞迴改進來實現,其中來自某一配置的經驗在外部契約下被評估,並用於構建其繼承者。協作則透過 LoRA 混合(Mixture-of-LoRA, MoL)架構來實現,該架構凍結基礎模型、組合專家 LoRA 適配器,並在每次使用者回合中選擇一個 LoRA。旗艦模型 Macaron-V1-Venti 結合了 744B GLM-5.2 基礎模型與四個分別用於聊天、智能體、程式碼和 GenUI 的 LoRA;基於 Qwen3.6 的 Macaron-V1-Tall(50B)則採用相同設計用於本地部署。本報告將 Macaron-V1 呈現為一個涵蓋架構、演算法和基礎設施的協同設計系統。MoL 架構透過可擴展的 LoRA 專家支援持續學習。演算法結合了模型-測試框架協同設計與遞迴自我改進迴圈,其中包括 UI4A 元件原生 GenUI 測試框架、有狀態動作基底、版本化 HCP 契約,以及智能體強化學習框架 MindForge。支援性基礎設施包括後期訓練平台 MinT、長上下文強化學習方法 LongStraw,以及針對稀疏 MoE 和 DSA 基礎模型的穩定性技術。我們在個人智能、GenUI 和一般能力基準上,對照前沿基線評估了 Macaron-V1。我們的結果驗證了當前系統,而持續學習和集體智能帶來的複合增益仍是開放問題。
English
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.