Institution profile

LexisNexis Risk Solutions

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

May 27, 2026

This study addresses the sharp performance degradation of existing linear probes in detecting deception by large language models under distribution shift, a phenomenon whose root cause remains unclear. Through a systematic analysis of the geometric structure of deceptive representations in the Gemma 3 model family, the work proposes a cross-domain transfer matrix, a permutation-null-baseline multidimensional probe, an entropy-residualization test, and an evaluation framework incorporating eight stylistic perturbations. The research demonstrates for the first time that probe fragility stems from narrow training distributions rather than inherent architectural limitations, thereby refuting prevailing hypotheses such as single-direction encoding, entropy-based proxies, and linear subspace assumptions. The proposed style-augmented probes achieve average AUROCs of 0.979–0.983 on unseen styles, while multidimensional probes (k≥5) successfully recover distributed weak signals, overturning the inverse scaling conjecture.

0 citationsRead paper

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

May 08, 2026

This work addresses the critical threat of backdoor attacks to language model security, which existing interpretability-based detection methods struggle to identify effectively. The study proposes a novel backdoor detection mechanism grounded in activation divergence, revealing for the first time that backdoor triggers manifest as directional shifts in activation patterns rather than sparse feature activations. Through systematic analysis of multi-layer representations in SmolLM2-360M under both LoRA and full-rank fine-tuning settings, the authors compare the backdoor identification capabilities of Crosscoders and differential sparse autoencoders (Diff-SAE). Experimental results demonstrate that Diff-SAE achieves a backdoor isolation score of 0.40, perfect precision of 1.0, and zero false positives across multiple configurations—substantially outperforming Crosscoders, which attain a backdoor isolation score below 0.02. These findings underscore the fundamental advantage of differential representations in isolating backdoor behaviors.

0 citationsRead paper

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

Apr 21, 2026

This study investigates the capacity of small language models to effectively use tools without relying on complex adaptation mechanisms. Focusing on Llama-3.2-3B-Instruct, the authors systematically evaluate four adaptation strategies—hypernetwork-generated LoRA weights, few-shot prompting, document-based prompting, and value-guided beam search—across four tool-use benchmarks. Experimental results demonstrate that few-shot prompting yields a 21.5% performance gain, document prompting contributes an additional 5.0%, while hypernetwork-generated LoRA weights show no significant improvement. Notably, the 3B-parameter model achieves 79.7% of GPT-5’s average performance at only one-tenth of the inference latency. These findings underscore the pivotal role of prompt engineering in enabling efficient tool use with lightweight models and offer a promising direction for resource-constrained settings.

0 citationsRead paper
Recent publications

Latest Papers

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations

May 27, 2026

This study addresses the sharp performance degradation of existing linear probes in detecting deception by large language models under distribution shift, a phenomenon whose root cause remains unclear. Through a systematic analysis of the geometric structure of deceptive representations in the Gemma 3 model family, the work proposes a cross-domain transfer matrix, a permutation-null-baseline multidimensional probe, an entropy-residualization test, and an evaluation framework incorporating eight stylistic perturbations. The research demonstrates for the first time that probe fragility stems from narrow training distributions rather than inherent architectural limitations, thereby refuting prevailing hypotheses such as single-direction encoding, entropy-based proxies, and linear subspace assumptions. The proposed style-augmented probes achieve average AUROCs of 0.979–0.983 on unseen styles, while multidimensional probes (k≥5) successfully recover distributed weak signals, overturning the inverse scaling conjecture.

0 citationsRead paper

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

May 08, 2026

This work addresses the critical threat of backdoor attacks to language model security, which existing interpretability-based detection methods struggle to identify effectively. The study proposes a novel backdoor detection mechanism grounded in activation divergence, revealing for the first time that backdoor triggers manifest as directional shifts in activation patterns rather than sparse feature activations. Through systematic analysis of multi-layer representations in SmolLM2-360M under both LoRA and full-rank fine-tuning settings, the authors compare the backdoor identification capabilities of Crosscoders and differential sparse autoencoders (Diff-SAE). Experimental results demonstrate that Diff-SAE achieves a backdoor isolation score of 0.40, perfect precision of 1.0, and zero false positives across multiple configurations—substantially outperforming Crosscoders, which attain a backdoor isolation score below 0.02. These findings underscore the fundamental advantage of differential representations in isolating backdoor behaviors.

0 citationsRead paper

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

Apr 21, 2026

This study investigates the capacity of small language models to effectively use tools without relying on complex adaptation mechanisms. Focusing on Llama-3.2-3B-Instruct, the authors systematically evaluate four adaptation strategies—hypernetwork-generated LoRA weights, few-shot prompting, document-based prompting, and value-guided beam search—across four tool-use benchmarks. Experimental results demonstrate that few-shot prompting yields a 21.5% performance gain, document prompting contributes an additional 5.0%, while hypernetwork-generated LoRA weights show no significant improvement. Notably, the 3B-parameter model achieves 79.7% of GPT-5’s average performance at only one-tenth of the inference latency. These findings underscore the pivotal role of prompt engineering in enabling efficient tool use with lightweight models and offer a promising direction for resource-constrained settings.

0 citationsRead paper