Praxist: From Experimental Artifacts to Solution Lineages
研究解决了自主研发代理在实验中难以追踪改进原因和成本过高的问题,通过引入Praxist系统,将可重复的实验结果转化为证据图谱,以更低的成本获得更优的结果。
研究解决了自主研发代理在实验中难以追踪改进原因和成本过高的问题,通过引入Praxist系统,将可重复的实验结果转化为证据图谱,以更低的成本获得更优的结果。
This work addresses the vulnerability of pretrained Transformers to out-of-distribution (OOD) inputs—such as misspellings or jailbreak prompts—which perturb internal representations and compromise model reliability and safety. Moving beyond input-space analyses, the study innovatively reframes the OOD problem within the model’s internal computation, employing sparse autoencoders to dissect activation patterns under OOD conditions. This reveals a marked proliferation of spurious or erroneous concepts in hidden representations. Leveraging this mechanistic insight, the authors propose an inference-time diagnostic method and a targeted fine-tuning strategy that operate directly on internal representations. Their approach effectively quantifies distributional shift in prompts and substantially enhances the robustness and safety of large language models against adversarial and anomalous inputs, establishing a novel paradigm for secure deployment.
研究解决了自主研发代理在实验中难以追踪改进原因和成本过高的问题,通过引入Praxist系统,将可重复的实验结果转化为证据图谱,以更低的成本获得更优的结果。
This work addresses the vulnerability of pretrained Transformers to out-of-distribution (OOD) inputs—such as misspellings or jailbreak prompts—which perturb internal representations and compromise model reliability and safety. Moving beyond input-space analyses, the study innovatively reframes the OOD problem within the model’s internal computation, employing sparse autoencoders to dissect activation patterns under OOD conditions. This reveals a marked proliferation of spurious or erroneous concepts in hidden representations. Leveraging this mechanistic insight, the authors propose an inference-time diagnostic method and a targeted fine-tuning strategy that operate directly on internal representations. Their approach effectively quantifies distributional shift in prompts and substantially enhances the robustness and safety of large language models against adversarial and anomalous inputs, establishing a novel paradigm for secure deployment.