Safin-1:透過記憶原生狀態演化實現由內而生的安全
Safin-1: Safety from Within through Memory-Native State Evolution
August 31, 2026
作者: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang, Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang, Ning Ding, Xia Hu, Bowen Zhou, Youbang Sun, Chaochao Lu
cs.AI
摘要
長時間跨度的複雜任務要求基礎模型在擴展的互動過程中累積資訊、維持內部狀態並進行適應。安全性應是模型本身的內在屬性,而非僅依賴外部防護措施或事後對齊(如監督式微調)的行為約束。這催生了「由內而生的安全性」(Safety from Within)理念,亦即透過模型自身的原生計算來表徵與調用與安全相關的能力。我們提出 Safin-1,一個透過記憶路由與狀態演化實現此原則的基礎模型系列。Safin-1 建構於 MARCH(Memory-Anchor Routing across Context History,跨上下文歷史的記憶錨點路由)架構之上,該網路架構能維護結構化的記憶狀態,並透過內容條件化路由選擇性地檢索相關歷史資訊。它支援在測試期間適應持續性的能力狀態,而無需反覆修改主幹網路,從而在共享基礎上實現受控的專門化。我們透過「安全狀態」(Safety State)在下游安全任務上探討此一介面,並展現了基於狀態的有效適應,帶來顯著的安全性提升。更廣泛而言,路由狀態介面將上下文記憶與持續性能力適應統一於模型的原生計算之中,將記憶從被動的先前行上下文記錄,重新詮釋為一個用於維持與演化模型行為的主動基質。整體能力、長上下文理解、檢索與效率等方面的評測進一步驗證了 Safin-1 的有效性。這些發現為「將安全性作為一種狀態原生且可持續適應的能力」提供了實現路徑。本工作僅是「由內而生的安全性」之初步架構探索,要實現此更宏大的願景,仍需大量後續研究。
English
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.