Safin-1: Safety from Within through Memory-Native State Evolution

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Safin-1模型,通过内存路由和状态演化实现内在安全性,解决长时复杂任务中模型安全性和适应性问题。
📝 Abstract
Long-horizon complex tasks require foundation models to accumulate information, maintain internal states, and adapt over extended interactions. Safety should be an intrinsic property of the model itself, rather than a behavioral constraint relying solely on external safeguards or post-hoc alignment such as supervised fine-tuning. This motivates Safety from Within, where safety-relevant capabilities are represented and invoked through the model's native computation. We present Safin-1, a family of foundation models realizing this principle through memory routing and state evolution. Safin-1 is built on Memory-Anchor Routing across Context History (MARCH), a network architecture that maintains structured memory states and selectively retrieves relevant historical information through content-conditioned routing. It supports test-time adaptation of persistent capability states without repeatedly modifying the backbone, enabling controlled specialization over a shared foundation. We investigate this interface on downstream safety tasks through a Safety State, demonstrating effective state-based adaptation with substantial safety improvements. More broadly, the routed-state interface unifies contextual memory and persistent capability adaptation within the model's native computation, reframing memory from a passive record of prior context into an active substrate for maintaining and evolving model behavior. Evaluations across general capabilities, long-context understanding, retrieval, and efficiency further validate Safin-1. These findings provide a path toward safety as a state-native and adaptively maintainable capability. This work is only an initial architectural exploration of Safety from Within, and substantial further work is needed to realize this broader vision.
Problem

Research questions and friction points this paper is trying to address.

Safety from Within
long-horizon complex tasks
foundation models
intrinsic property
memory-native state evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety from Within
memory-native state evolution
MARCH (Memory-Anchor Routing across Context History)
test-time adaptation
state-based adaptation
M
Ming Zhang
Shanghai AI Laboratory
K
Kaisen Yang
Shanghai AI Laboratory
S
Shu Yu
Shanghai AI Laboratory
Ermo Hua
Ermo Hua
Tsinghua University
Physics-driven Foundation Model
Z
Zhekai Chen
Shanghai AI Laboratory
C
Cheng Jin
Shanghai AI Laboratory
J
Jingnan Zheng
Shanghai AI Laboratory
Y
Yi Zhang
Shanghai AI Laboratory
Z
Zhongtian Ma
Shanghai AI Laboratory
J
Jiawei Zhou
Shanghai AI Laboratory
Sirui Chen
Sirui Chen
University of Illinois Urbana-Champaign
Reinforcement LearningInformation Retrieval
Qiaosheng Zhang
Qiaosheng Zhang
Department of Anesthesiology, New York University School of Medcine
NeuroscienceNeural EngineeringPain CircuitryBrain Machine InterfaceNeural Signal Processing
X
Xiang Wang
Shanghai AI Laboratory
N
Ning Ding
Shanghai AI Laboratory
Xia Hu
Xia Hu
Google DeepMind
Deep LearningMachine LearningMultimodal
B
Bowen Zhou
Shanghai AI Laboratory
Youbang Sun
Youbang Sun
Assistant Researcher, Tsinghua University; Northeastern University; Texas A&M University
Distributed OptimizationMulti-Agent RLRiemannian OptimizationFederated Learning
Chaochao Lu
Chaochao Lu
Shanghai AI Laboratory
Causal AI