Safin-1: 메모리-네이티브 상태 진화를 통한 내재적 안전성
Safin-1: Safety from Within through Memory-Native State Evolution
August 31, 2026
저자: Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang, Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang, Ning Ding, Xia Hu, Bowen Zhou, Youbang Sun, Chaochao Lu
cs.AI
초록
장기적인 복잡한 과제(Long-horizon complex tasks)는 파운데이션 모델이 정보를 축적하고, 내부 상태를 유지하며, 확장된 상호작용에 걸쳐 적응할 것을 요구한다. 안전성은 지도 미세조정과 같은 외부 안전장치나 사후 정렬에만 의존하는 행동적 제약이 아니라, 모델 자체의 고유한 속성이어야 한다. 이러한 동기는 안전 관련 능력이 모델의 고유 계산을 통해 표현되고 호출되는 '내재적 안전성(Safety from Within)'을 추구하게 한다. 본 논문은 메모리 라우팅과 상태 진화를 통해 이 원리를 구현한 파운데이션 모델 제품군인 Safin-1을 제시한다. Safin-1은 컨텍스트 이력 전반의 메모리 앵커 라우팅(MARCH: Memory-Anchor Routing across Context History)을 기반으로 구축되었으며, 이는 구조화된 메모리 상태를 유지하고 내용 조건부 라우팅을 통해 관련 과거 정보를 선택적으로 검색하는 네트워크 아키텍처이다. 또한 백본을 반복적으로 수정하지 않고도 지속적 능력 상태의 테스트 시간 적응을 지원하여, 공유 파운데이션 위에서 제어된 특수화를 가능하게 한다. 본 연구는 안전 상태(Safety State)를 통해 다운스트림 안전 과제에서 이러한 인터페이스를 조사하였으며, 상태 기반 적응이 상당한 안전성 개선을 가져옴을 입증하였다. 더 넓게 보면, 라우팅된 상태 인터페이스는 맥락적 메모리와 지속적 능력 적응을 모델의 고유 계산 내에서 통합함으로써, 메모리를 이전 컨텍스트의 수동적 기록에서 모델 행동을 유지하고 진화시키는 능동적 기반으로 재구성한다. 일반 능력, 장기 컨텍스트 이해, 검색 및 효율성에 대한 평가는 Safin-1을 추가로 검증한다. 이러한 발견은 안전성을 상태 고유적이고 적응적으로 유지 가능한 능력으로 구현하는 경로를 제시한다. 본 연구는 내재적 안전성의 광범위한 비전을 실현하기 위한 초기 아키텍처 탐색에 불과하며, 이를 완성하기 위해서는 상당한 추가 연구가 필요하다.
English
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.