Institution profile

Sapient Intelligence

Industry researchnorthamerica · us
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

Jun 24, 2026

This work addresses the vulnerability of pretrained Transformers to out-of-distribution (OOD) inputs—such as misspellings or jailbreak prompts—which perturb internal representations and compromise model reliability and safety. Moving beyond input-space analyses, the study innovatively reframes the OOD problem within the model’s internal computation, employing sparse autoencoders to dissect activation patterns under OOD conditions. This reveals a marked proliferation of spurious or erroneous concepts in hidden representations. Leveraging this mechanistic insight, the authors propose an inference-time diagnostic method and a targeted fine-tuning strategy that operate directly on internal representations. Their approach effectively quantifies distributional shift in prompts and substantially enhances the robustness and safety of large language models against adversarial and anomalous inputs, establishing a novel paradigm for secure deployment.

0 citationsRead paper
Recent publications

Latest Papers

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

Jun 24, 2026

This work addresses the vulnerability of pretrained Transformers to out-of-distribution (OOD) inputs—such as misspellings or jailbreak prompts—which perturb internal representations and compromise model reliability and safety. Moving beyond input-space analyses, the study innovatively reframes the OOD problem within the model’s internal computation, employing sparse autoencoders to dissect activation patterns under OOD conditions. This reveals a marked proliferation of spurious or erroneous concepts in hidden representations. Leveraging this mechanistic insight, the authors propose an inference-time diagnostic method and a targeted fine-tuning strategy that operate directly on internal representations. Their approach effectively quantifies distributional shift in prompts and substantially enhances the robustness and safety of large language models against adversarial and anomalous inputs, establishing a novel paradigm for secure deployment.

0 citationsRead paper