ChatPaper.aiChatPaper

Safin-1:メモリネイティブ状態進化による内部からの安全性

Safin-1: Safety from Within through Memory-Native State Evolution

August 31, 2026
著者: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang, Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang, Ning Ding, Xia Hu, Bowen Zhou, Youbang Sun, Chaochao Lu
cs.AI

要旨

長期的な複雑なタスクでは、基盤モデルが情報を蓄積し、内部状態を維持し、長期間の相互作用にわたって適応することが求められる。安全性は、外部の保護手段や教師ありファインチューニングなどの事後的なアラインメントのみに依存する行動制約ではなく、モデル自体の内在的特性であるべきである。この考えに動機づけられたSafety from Within(内部からの安全性)では、安全に関わる能力がモデル本来の計算を通じて表現され、呼び出される。本稿では、メモリルーティングと状態進化によりこの原理を実現する基盤モデル群Safin-1を提案する。Safin-1は、文脈履歴にわたるメモリアンカールーティング(Memory-Anchor Routing across Context History: MARCH)に基づく。MARCHは、構造化されたメモリ状態を維持し、内容に応じて条件付けられたルーティングにより関連する履歴情報を選択的に取得するネットワークアーキテクチャである。これにより、バックボーンを繰り返し変更することなく永続的な能力状態のテスト時適応が可能となり、共有された基盤上での制御された専門化を実現する。本稿では、このインターフェースを下流の安全性タスクにおいてセーフティステート(Safety State)を通じて調査し、状態ベースの適応が効果的であり、安全性を大幅に向上させることを示す。さらに広く見れば、ルーティングされた状態インターフェースは、モデル固有の計算内で文脈メモリと永続的な能力適応を統合し、メモリを過去の文脈の受動的な記録から、モデルの挙動を維持し進化させる能動的な基盤へと再定義する。汎用能力、長文脈理解、検索、効率性にわたる評価により、Safin-1の有効性がさらに検証される。これらの知見は、安全性を状態内在的かつ適応的に維持可能な能力として扱うための道筋を提供する。本研究はSafety from Withinの初期のアーキテクチャ探求に過ぎず、このより広範なビジョンを実現するには、さらなる大規模な研究が必要である。
English
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.