Institution profile

LASR Labs

Industry researchnorthamerica · us
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Building Better Deception Probes Using Targeted Instruction Pairs

Feb 01, 2026

This work addresses the vulnerability of existing linear probes to spurious correlations when detecting AI deception, which often leads to false positives on non-deceptive responses. To mitigate this, the authors propose a targeted instruction-pair design method grounded in a taxonomy of deceptive behaviors. By constructing interpretable instruction pairs that isolate specific deception types, the approach trains linear probes to focus on deceptive intent rather than superficial content patterns. Experimental results demonstrate that instruction selection is the dominant factor in probe performance, accounting for 70.6% of variance. Probes tailored to specific threat models significantly outperform general-purpose detectors, achieving higher detection accuracy and lower false-positive rates on evaluation datasets. This study underscores the importance of aligning probing mechanisms with concrete deception categories and establishes a new paradigm for interpretable AI safety evaluation.

0 citationsRead paper

CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring

May 29, 2025

Autonomous deployment of large language models (LLMs) poses security risks, as malicious actors may stealthily induce harmful behaviors. Method: We propose a novel safety supervision paradigm based on monitoring intermediate chain-of-thought (CoT) reasoning—departing from conventional output-only monitoring. Our approach introduces a hybrid monitoring protocol that independently evaluates both the CoT process and the final output, followed by weighted fusion for dual-path collaborative decision-making. Contribution/Results: We are the first to identify and systematically quantify how CoT reasoning can be adversarially rationalized to evade detection, establishing its failure boundaries. We further develop a red-teaming evaluation framework and a cross-model robustness assessment suite. Experiments demonstrate over a 4× improvement in detection rate for subtle deception scenarios; across multiple models and tasks, our method consistently outperforms single-path baselines, achieving up to a 27-percentage-point gain in accuracy.

0 citationsRead paper
Recent publications

Latest Papers

Building Better Deception Probes Using Targeted Instruction Pairs

Feb 01, 2026

This work addresses the vulnerability of existing linear probes to spurious correlations when detecting AI deception, which often leads to false positives on non-deceptive responses. To mitigate this, the authors propose a targeted instruction-pair design method grounded in a taxonomy of deceptive behaviors. By constructing interpretable instruction pairs that isolate specific deception types, the approach trains linear probes to focus on deceptive intent rather than superficial content patterns. Experimental results demonstrate that instruction selection is the dominant factor in probe performance, accounting for 70.6% of variance. Probes tailored to specific threat models significantly outperform general-purpose detectors, achieving higher detection accuracy and lower false-positive rates on evaluation datasets. This study underscores the importance of aligning probing mechanisms with concrete deception categories and establishes a new paradigm for interpretable AI safety evaluation.

0 citationsRead paper

CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring

May 29, 2025

Autonomous deployment of large language models (LLMs) poses security risks, as malicious actors may stealthily induce harmful behaviors. Method: We propose a novel safety supervision paradigm based on monitoring intermediate chain-of-thought (CoT) reasoning—departing from conventional output-only monitoring. Our approach introduces a hybrid monitoring protocol that independently evaluates both the CoT process and the final output, followed by weighted fusion for dual-path collaborative decision-making. Contribution/Results: We are the first to identify and systematically quantify how CoT reasoning can be adversarially rationalized to evade detection, establishing its failure boundaries. We further develop a red-teaming evaluation framework and a cross-model robustness assessment suite. Experiments demonstrate over a 4× improvement in detection rate for subtle deception scenarios; across multiple models and tasks, our method consistently outperforms single-path baselines, achieving up to a 27-percentage-point gain in accuracy.

0 citationsRead paper