Institution profile

Lakera AI

Industry researcheurope · ch
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Identifying Causal Effects Using a Single Proxy Variable

Apr 10, 2026

This study addresses the challenge of identifying causal effects in the presence of unobserved confounders by introducing the SPICE condition, which ensures identifiability under the assumption that only a single observed proxy variable is available and its generative mechanism is known. Building on this condition, the authors develop SPICE-Net, a general-purpose neural network framework capable of handling both discrete and continuous treatment variables. The work substantially extends existing proxy-based identification theory to high-dimensional settings, nonlinear functional relationships, and broader distributional classes. It presents the first learnable, end-to-end approach to causal identification and estimation grounded in the completeness assumption, provides rigorous theoretical proof of identifiability under the SPICE condition, and empirically demonstrates the method’s effectiveness across diverse treatment types.

0 citationsRead paper

Invariance-Based Dynamic Regret Minimization

Mar 04, 2026

This work addresses the challenge in stochastic non-stationary linear bandits where time-varying parameters cause conventional methods to prematurely discard historical data, thereby losing valuable information. To mitigate this issue, the authors propose decomposing the reward model into stationary and non-stationary components and introduce invariance modeling—leveraging stable structures within historical data to reduce the effective problem dimensionality—within a dynamic regret minimization framework. The proposed ISD-linUCB algorithm integrates contextual linear modeling, dynamic weighting, and invariance identification to enable efficient online decision-making in non-stationary environments. Both theoretical analysis and empirical experiments demonstrate that, particularly in rapidly changing settings with sufficient historical data, the method achieves significantly lower dynamic regret compared to existing approaches.

0 citationsRead paper

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

Jan 21, 2026

This work proposes a novel generalized method of moments (GMM) framework that treats multiple experimental environments as high-dimensional instrumental variables to address causal effect estimation under challenging conditions: covariates and outcomes are not jointly observed, unobserved confounders are present, and each environment contains only a very small sample size. By integrating cross-fitting sample splitting, ℓ₁ regularization, and post-regularization refitting, the proposed estimator achieves consistent causal effect estimation even as the number of environments grows to infinity while the sample size per environment remains fixed. Moreover, it effectively identifies sparse causal structures, thereby overcoming the inconsistency inherent in conventional two-sample instrumental variable approaches under this setting.

0 citationsRead paper

A Safety and Security Framework for Real-World Agentic Systems

Nov 26, 2025

This paper addresses novel safety and security risks—such as tool misuse, cascading action chains, and unintended control amplification—that arise in enterprise-grade agentic AI systems due to dynamic interactions among models, coordinators, tools, and data during real-world deployment. Methodologically, it establishes the first unified, dynamic safety-and-security framework for agentic systems, featuring a fine-grained risk taxonomy and a mechanistic analysis of safety-security coupling in dynamic execution. It introduces a closed-loop governance paradigm integrating auxiliary AI-enhanced risk perception, collaborative sandboxed execution, and AI-driven red-teaming. Evaluated end-to-end on the NVIDIA AI-Q Research Assistant platform, the framework identifies and mitigates over ten classes of emergent risks. Additionally, it releases an open-source benchmark dataset comprising 10,000+ adversarial and defensive agent trajectories, providing both theoretical foundations and empirical infrastructure for agentic AI safety research.

0 citationsRead paper

Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents

Oct 26, 2025

Existing LLM safety evaluation frameworks lack systematic characterization of security propagation risks arising when LLMs serve as the backbone of AI agents. To address this, we propose the “Threat Snapshot” method, which isolates critical execution states wherein LLM vulnerabilities manifest during agent operation, enabling precise identification and categorization of security risks. Based on this methodology, we introduce b³—the first agent-centric safety benchmark—comprising 194,331 crowdsourced adversarial examples, and empirically evaluate 31 mainstream LLMs. Our findings reveal that enhanced reasoning capabilities correlate with improved safety, whereas model scale exhibits no statistically significant relationship with security performance. We publicly release the dataset, evaluation code, and benchmark infrastructure to establish a reproducible, quantifiable, and scalable assessment paradigm for LLM safety design.

0 citationsRead paper
Recent publications

Latest Papers

Identifying Causal Effects Using a Single Proxy Variable

Apr 10, 2026

This study addresses the challenge of identifying causal effects in the presence of unobserved confounders by introducing the SPICE condition, which ensures identifiability under the assumption that only a single observed proxy variable is available and its generative mechanism is known. Building on this condition, the authors develop SPICE-Net, a general-purpose neural network framework capable of handling both discrete and continuous treatment variables. The work substantially extends existing proxy-based identification theory to high-dimensional settings, nonlinear functional relationships, and broader distributional classes. It presents the first learnable, end-to-end approach to causal identification and estimation grounded in the completeness assumption, provides rigorous theoretical proof of identifiability under the SPICE condition, and empirically demonstrates the method’s effectiveness across diverse treatment types.

0 citationsRead paper

Invariance-Based Dynamic Regret Minimization

Mar 04, 2026

This work addresses the challenge in stochastic non-stationary linear bandits where time-varying parameters cause conventional methods to prematurely discard historical data, thereby losing valuable information. To mitigate this issue, the authors propose decomposing the reward model into stationary and non-stationary components and introduce invariance modeling—leveraging stable structures within historical data to reduce the effective problem dimensionality—within a dynamic regret minimization framework. The proposed ISD-linUCB algorithm integrates contextual linear modeling, dynamic weighting, and invariance identification to enable efficient online decision-making in non-stationary environments. Both theoretical analysis and empirical experiments demonstrate that, particularly in rapidly changing settings with sufficient historical data, the method achieves significantly lower dynamic regret compared to existing approaches.

0 citationsRead paper

Many Experiments, Few Repetitions, Unpaired Data, and Sparse Effects: Is Causal Inference Possible?

Jan 21, 2026

This work proposes a novel generalized method of moments (GMM) framework that treats multiple experimental environments as high-dimensional instrumental variables to address causal effect estimation under challenging conditions: covariates and outcomes are not jointly observed, unobserved confounders are present, and each environment contains only a very small sample size. By integrating cross-fitting sample splitting, ℓ₁ regularization, and post-regularization refitting, the proposed estimator achieves consistent causal effect estimation even as the number of environments grows to infinity while the sample size per environment remains fixed. Moreover, it effectively identifies sparse causal structures, thereby overcoming the inconsistency inherent in conventional two-sample instrumental variable approaches under this setting.

0 citationsRead paper

A Safety and Security Framework for Real-World Agentic Systems

Nov 26, 2025

This paper addresses novel safety and security risks—such as tool misuse, cascading action chains, and unintended control amplification—that arise in enterprise-grade agentic AI systems due to dynamic interactions among models, coordinators, tools, and data during real-world deployment. Methodologically, it establishes the first unified, dynamic safety-and-security framework for agentic systems, featuring a fine-grained risk taxonomy and a mechanistic analysis of safety-security coupling in dynamic execution. It introduces a closed-loop governance paradigm integrating auxiliary AI-enhanced risk perception, collaborative sandboxed execution, and AI-driven red-teaming. Evaluated end-to-end on the NVIDIA AI-Q Research Assistant platform, the framework identifies and mitigates over ten classes of emergent risks. Additionally, it releases an open-source benchmark dataset comprising 10,000+ adversarial and defensive agent trajectories, providing both theoretical foundations and empirical infrastructure for agentic AI safety research.

0 citationsRead paper

Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents

Oct 26, 2025

Existing LLM safety evaluation frameworks lack systematic characterization of security propagation risks arising when LLMs serve as the backbone of AI agents. To address this, we propose the “Threat Snapshot” method, which isolates critical execution states wherein LLM vulnerabilities manifest during agent operation, enabling precise identification and categorization of security risks. Based on this methodology, we introduce b³—the first agent-centric safety benchmark—comprising 194,331 crowdsourced adversarial examples, and empirically evaluate 31 mainstream LLMs. Our findings reveal that enhanced reasoning capabilities correlate with improved safety, whereas model scale exhibits no statistically significant relationship with security performance. We publicly release the dataset, evaluation code, and benchmark infrastructure to establish a reproducible, quantifiable, and scalable assessment paradigm for LLM safety design.

0 citationsRead paper