Institution profile

Korea Institute of Energy Technology

Academic institutionasia · kr
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

Jul 23, 2026

Existing long-sequence memory models struggle to simultaneously achieve lossless long-term retention and effective overwriting of outdated information. This work proposes Naju, the first model to decouple forgetting and writing mechanisms within a discrete state space. By employing a learnable sigmoid forget gate, an independent write gate, and input-dependent linear read-write mappings, Naju decomposes recurrent updates into orthogonal operations. This design overcomes the theoretical trade-off between retention rate and write strength inherent in conventional single-gate architectures, all while preserving linear time and space complexity. Experiments demonstrate that Naju maintains superior memory retention and overwrite capabilities even when trained on sequences four times longer than baseline lengths, outperforming Mamba on WikiText-103, the Long Range Arena benchmark, and multi-query associative recall tasks, with performance comparable to Transformers.

0 citationsRead paper
Recent publications

Latest Papers

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

Jul 23, 2026

Existing long-sequence memory models struggle to simultaneously achieve lossless long-term retention and effective overwriting of outdated information. This work proposes Naju, the first model to decouple forgetting and writing mechanisms within a discrete state space. By employing a learnable sigmoid forget gate, an independent write gate, and input-dependent linear read-write mappings, Naju decomposes recurrent updates into orthogonal operations. This design overcomes the theoretical trade-off between retention rate and write strength inherent in conventional single-gate architectures, all while preserving linear time and space complexity. Experiments demonstrate that Naju maintains superior memory retention and overwrite capabilities even when trained on sequences four times longer than baseline lengths, outperforming Mamba on WikiText-103, the Long Range Arena benchmark, and multi-query associative recall tasks, with performance comparable to Transformers.

0 citationsRead paper