ChatPaper.aiChatPaper

Macaron-V1:自己改善とMixture-of-LoRAによるオープン継続学習へ向けて

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

August 10, 2026
著者: Mind Lab, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong, Kuss Koo, Jaron Lee, Andrew Lei, Alexy Li, Dawn Li, Lucian Li, Ray Li, Ricardo Li, Smith Li, Theo Li, Allen Lin, Elliot Lin, Fan Lin, Chen Ling, Kairus Liu, Kieran Liu, Logan Liu, Neo Liu, Xiang Liu, Yuxin Lu, Maeve Luo, Pony Ma, Verity Niu, Cole Qiao, Guian Qiu, Vince Qu, Sentry, Niko Song, Vincent Wang, Bo Wu, Rio Yang, Evelyn Ye, Fiona Ye, Ina Ye, Regis Ye, Josh Ying, Atlas Zeng, Danney Zeng, Salmon Zhan, Anya Zhang, Di Zhang, Mia Zhang, Sueky Zhang, Wei Zhao, Ada Zhou, Adrian Zhou, Yuhua Zhou, Juno Zhu, Murphy Zhuang
cs.AI

要旨

Macaron-V1は、経験的知能のためのオープンなエージェントモデルファミリーである。すなわち、実環境での経験から学習し、デプロイ後も学習を継続するものである。本ファミリーは、2つのシステム目標を中心に設計されている。適応(Adaptation)は、バージョン管理されたモデル-ハーネス対の再帰的改善を通じて追求される。ある構成から得られた経験は、外部契約の下で評価され、その後継の構築に用いられる。協調(Collaboration)は、Mixture-of-LoRA(MoL)アーキテクチャによって実現される。このアーキテクチャは、ベースモデルを凍結し、専門化されたLoRAアダプターを構成し、ユーザーターンごとに1つのLoRAを選択する。フラグシップモデルであるMacaron-V1-Ventiは、744BのGLM-5.2ベースに、チャット、エージェント、コーディング、GenUI向けの4つのLoRAを組み合わせている。Qwen3.6ベースのMacaron-V1-Tall(50B)は、ローカルデプロイメント向けに同じ設計を採用している。 本レポートでは、Macaron-V1を、アーキテクチャ、アルゴリズム、インフラストラクチャにわたる共設計システムとして提示する。MoLアーキテクチャは、拡張可能なLoRAスペシャリストを通じて継続学習を支援する。アルゴリズムは、モデル-ハーネス共設計(Model-Harness Co-design)と再帰的自己改善ループを組み合わせたものであり、具体的には、UI4AコンポーネントネイティブGenUIハーネス、状態を保持するアクション基盤、バージョン管理されたHCP契約、エージェント型RLフレームワークMindForgeを含む。これを支援するインフラストラクチャには、ポストトレーニングプラットフォームMinT、長文脈RL手法LongStraw、スパースMoEおよびDSAベースモデル向けの安定性技術が含まれる。我々は、Macaron-V1を、パーソナル知能、GenUI、および一般的な能力ベンチマークにおいて、最先端のベースラインと比較して評価する。我々の結果は現行システムの有効性を裏付けるものである。一方、継続学習と集合知による複利的な利得の実現は、依然として未解決の課題である。
English
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.