Institution profile

AI Security Institute

Academic institutionnorthamerica · us
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

Jun 10, 2026

Current language models lack a reliable evaluation framework for deceptive behavior due to the inaccessibility of their true beliefs. This work introduces the first set of reasoning-based model organisms (13 in total) with verifiable internal beliefs and constructs Varied Deception, a prompt-based lying benchmark. Coupled with a novel probe-training method, Did-You-Lie (DYL), the study systematically evaluates four deception detection approaches across models ranging from 2B to 1T parameters. Experiments reveal that while all detectors improve with model capability on prompt-induced lying tasks, only chain-of-thought–based discriminators maintain high accuracy (balanced accuracy of 0.82) on trained model organisms; other methods degrade significantly. This work establishes the first reliable methodology for assessing inconsistencies between a model’s internal beliefs and its generated outputs.

0 citationsRead paper

RealityTest: How People Probe AI Identity and Whether Models Disclose It

May 29, 2026

Existing evaluations of AI identity disclosure are largely confined to English, synthetic queries, and text-only settings, failing to capture real-world identity probing behaviors in authentic conversations. This work proposes RealityTest—the first large-scale, multilingual, multimodal benchmark for assessing AI identity disclosure grounded in real human interactions—drawing on data from approximately 750 participants across 49 countries. The study systematically evaluates 17 text-based and 6 speech-based models, revealing that only 31% of users directly inquire about an AI’s identity in ambiguous contexts, and that real-world questions exhibit substantially greater diversity than synthetic ones. Notably, a single suppression instruction reduces disclosure rates below 30%, exposing significant biases in current safety assessments due to narrow evaluation data. This work provides the first empirical evidence of the critical influence of question phrasing and contextual framing on AI identity disclosure.

0 citationsRead paper

Automated alignment is harder than you think

May 07, 2026

This study addresses a critical vulnerability in automated alignment research: in ambiguous tasks, reliance on human supervision—constrained by cognitive limitations and ill-defined evaluation criteria—can lead to subtle but consequential errors, potentially resulting in false assurances of AI safety and the deployment of misaligned artificial superintelligence. The work presents the first systematic argument that AI agents are more prone than humans to generate misleading conclusions in such settings, with this risk amplified by four interrelated factors: optimization pressure, disparities in error types, the infeasibility of evaluating certain arguments, and output correlations. Integrating insights from alignment theory, scalable oversight, and generalization analysis, the paper exposes significant pitfalls in current automated alignment paradigms and underscores the urgent need for AI systems capable of reliably handling ambiguity, while highlighting novel challenges for generalization and supervision methodologies.

0 citationsRead paper

Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

Mar 11, 2026

This study systematically evaluates the autonomous offensive capabilities of state-of-the-art AI models in complex, multi-stage cyberattacks involving heterogeneous action sequences. To this end, two specialized network ranges—a 32-step enterprise network and a 7-step industrial control system attack scenario—were constructed, and seven large language models released over an 18-month period were tested under varying inference compute budgets. The work introduces the first automated evaluation framework based on multi-step attack ranges, integrating heterogeneous attack action modeling with high-compute inference scheduling to quantitatively disentangle the independent effects of model generation advances and inference compute on long-chain attack performance. Results show that model performance scales log-linearly with inference compute: the latest models achieve an average of 9.8 out of 32 steps (peaking at 22) under a 10M-token budget and demonstrate, for the first time, partially stable execution in industrial control scenarios (averaging 1.2–1.4 out of 7 steps).

0 citationsRead paper

A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios

Feb 25, 2026

This study evaluates the risk of large language models (LLMs) being misused in sophisticated cybercrimes such as romance scams, CEO impersonation, and identity theft. To this end, we introduce the first reproducible, multi-turn interactive evaluation framework co-designed with law enforcement and policy experts. The framework decomposes malicious intent into seemingly benign queries to assess models’ ability to generate actionable information, benchmarking against standard web search and open-source LLMs with safety safeguards removed. Our findings indicate that mainstream closed-source models offer limited assistance for high-level criminal activities; however, their risk increases substantially when safety constraints are disabled. Moreover, multi-turn indirect requests prove more effective than explicit malicious prompts at bypassing current defenses, exposing critical limitations in existing safety strategies under complex adversarial scenarios.

0 citationsRead paper
Recent publications

Latest Papers

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms

Jun 10, 2026

Current language models lack a reliable evaluation framework for deceptive behavior due to the inaccessibility of their true beliefs. This work introduces the first set of reasoning-based model organisms (13 in total) with verifiable internal beliefs and constructs Varied Deception, a prompt-based lying benchmark. Coupled with a novel probe-training method, Did-You-Lie (DYL), the study systematically evaluates four deception detection approaches across models ranging from 2B to 1T parameters. Experiments reveal that while all detectors improve with model capability on prompt-induced lying tasks, only chain-of-thought–based discriminators maintain high accuracy (balanced accuracy of 0.82) on trained model organisms; other methods degrade significantly. This work establishes the first reliable methodology for assessing inconsistencies between a model’s internal beliefs and its generated outputs.

0 citationsRead paper

RealityTest: How People Probe AI Identity and Whether Models Disclose It

May 29, 2026

Existing evaluations of AI identity disclosure are largely confined to English, synthetic queries, and text-only settings, failing to capture real-world identity probing behaviors in authentic conversations. This work proposes RealityTest—the first large-scale, multilingual, multimodal benchmark for assessing AI identity disclosure grounded in real human interactions—drawing on data from approximately 750 participants across 49 countries. The study systematically evaluates 17 text-based and 6 speech-based models, revealing that only 31% of users directly inquire about an AI’s identity in ambiguous contexts, and that real-world questions exhibit substantially greater diversity than synthetic ones. Notably, a single suppression instruction reduces disclosure rates below 30%, exposing significant biases in current safety assessments due to narrow evaluation data. This work provides the first empirical evidence of the critical influence of question phrasing and contextual framing on AI identity disclosure.

0 citationsRead paper

Automated alignment is harder than you think

May 07, 2026

This study addresses a critical vulnerability in automated alignment research: in ambiguous tasks, reliance on human supervision—constrained by cognitive limitations and ill-defined evaluation criteria—can lead to subtle but consequential errors, potentially resulting in false assurances of AI safety and the deployment of misaligned artificial superintelligence. The work presents the first systematic argument that AI agents are more prone than humans to generate misleading conclusions in such settings, with this risk amplified by four interrelated factors: optimization pressure, disparities in error types, the infeasibility of evaluating certain arguments, and output correlations. Integrating insights from alignment theory, scalable oversight, and generalization analysis, the paper exposes significant pitfalls in current automated alignment paradigms and underscores the urgent need for AI systems capable of reliably handling ambiguity, while highlighting novel challenges for generalization and supervision methodologies.

0 citationsRead paper

Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

Mar 11, 2026

This study systematically evaluates the autonomous offensive capabilities of state-of-the-art AI models in complex, multi-stage cyberattacks involving heterogeneous action sequences. To this end, two specialized network ranges—a 32-step enterprise network and a 7-step industrial control system attack scenario—were constructed, and seven large language models released over an 18-month period were tested under varying inference compute budgets. The work introduces the first automated evaluation framework based on multi-step attack ranges, integrating heterogeneous attack action modeling with high-compute inference scheduling to quantitatively disentangle the independent effects of model generation advances and inference compute on long-chain attack performance. Results show that model performance scales log-linearly with inference compute: the latest models achieve an average of 9.8 out of 32 steps (peaking at 22) under a 10M-token budget and demonstrate, for the first time, partially stable execution in industrial control scenarios (averaging 1.2–1.4 out of 7 steps).

0 citationsRead paper

A Multi-Turn Framework for Evaluating AI Misuse in Fraud and Cybercrime Scenarios

Feb 25, 2026

This study evaluates the risk of large language models (LLMs) being misused in sophisticated cybercrimes such as romance scams, CEO impersonation, and identity theft. To this end, we introduce the first reproducible, multi-turn interactive evaluation framework co-designed with law enforcement and policy experts. The framework decomposes malicious intent into seemingly benign queries to assess models’ ability to generate actionable information, benchmarking against standard web search and open-source LLMs with safety safeguards removed. Our findings indicate that mainstream closed-source models offer limited assistance for high-level criminal activities; however, their risk increases substantially when safety constraints are disabled. Moreover, multi-turn indirect requests prove more effective than explicit malicious prompts at bypassing current defenses, exposing critical limitations in existing safety strategies under complex adversarial scenarios.

0 citationsRead paper