Safin-1:通过记忆原生状态演化的内部安全保障
Safin-1: Safety from Within through Memory-Native State Evolution
August 31, 2026
作者: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang, Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang, Ning Ding, Xia Hu, Bowen Zhou, Youbang Sun, Chaochao Lu
cs.AI
摘要
长时程复杂任务要求基础模型在长期交互中持续积累信息、维持内部状态并不断适应。安全性应当是模型自身的内在属性,而非仅仅依赖外部防护或事后对齐(如监督微调)等外部手段加以约束的行为。由此我们提出“由内而生的安全”(Safety from Within)理念:与安全相关的能力应在模型的原生计算中得到表征并被调用。我们基于这一理念构建了 Safin-1——一个通过记忆路由与状态演化实现该原则的基础模型系列。Safin-1 建立在跨上下文历史的记忆锚定路由(Memory-Anchor Routing across Context History, MARCH)架构之上。MARCH 维护结构化的记忆状态,并通过内容条件路由有选择地检索相关历史信息;同时,它支持在不反复修改主干网络的前提下对持久能力状态进行测试时自适应,从而在共享基础之上实现受控特化。我们以安全状态(Safety State)为接口,在下游安全任务上考察了这一机制,结果表明这种基于状态的自适应方式能够带来显著的安全性能提升。更广泛地看,路由状态接口将上下文记忆与持久能力自适应统一于模型的原生计算之中,使记忆从对先前上下文的被动记录,转变为维持并演化模型行为的主动基质。通用能力、长上下文理解、检索与效率等方面的评测进一步验证了 Safin-1。这些发现为将安全性塑造为一种状态原生、可自适应维护的能力指明了路径。需要指出的是,本工作仅是对“由内而生的安全”理念的初步架构探索,要实现这一更宏大的愿景仍需大量后续研究。
English
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.