Institution profile

Devoteam

Industry researcheurope · fr
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning

Jan 02, 2026arXiv.org

This work proposes a training-free method to assess the validity of mathematical reasoning by interpreting the attention matrices of large language models as dynamic graph adjacency matrices and applying spectral graph analysis to extract interpretable features—such as the Fiedler value, high-frequency energy ratio, graph signal smoothness, and spectral entropy. Evaluated across seven mainstream models, the approach achieves accuracy rates of 85.0%–95.6% (Cohen’s d = 3.30), which improve to 93%–95% after calibration. Notably, it correctly identifies valid proofs erroneously rejected by formal verifiers and reveals a shift in discriminative signals within Mistral-7B—from high-frequency energy ratio toward signal smoothness—highlighting the critical influence of attention mechanism design on reasoning reliability.

1 citationsRead paper

A Probe Direction Is a Property of Its Prompt

Aug 13, 2026

This study addresses a critical oversight in existing probing methods for assessing whether language models are aware of being evaluated: the decisive influence of prompt selection on measurement outcomes, which undermines cross-model comparability. Treating prompts as a core component of measurement design, the authors fix task content while systematically varying prompts and employ controlled experiments, probe direction analysis, variance decomposition, and surface-form ablation tests to quantify the contributions of prompts, models, and their interaction to observed scores. Findings reveal that models account for only a small fraction of score variance; prompt choice can reverse apparent scaling trends; and surface-level prompt features alone suffice to reproduce most published results. The work demonstrates that single-prompt probing yields unreliable comparisons and specifies the minimum number of prompts required for robust evaluation.

0 citationsRead paper

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

Aug 13, 2026

This study addresses the lack of a standardized measurement location for latent feature importance in sparse autoencoder (SAE) evaluation, which distorts comparisons across dictionaries. Through controlled experiments, we demonstrate that variation in measurement location accounts for up to 11.9% of evaluation variance—a discrepancy that intensifies with larger corpora. To resolve this, we propose a causal evaluation framework based on ablation, which aligns latent features via decoder similarity, and introduce a simple yet effective standardization protocol: fixing the measurement location with a single line of code. This protocol disentangles dictionary performance from positional bias, and its necessity and efficacy are validated through multiple controlled experiments and audits of published works, substantially improving the consistency and comparability of SAE evaluations.

0 citationsRead paper

Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology

Feb 08, 2026

This work addresses the challenge of hallucination in autonomous agents deployed in open environments when using external tools, where existing approaches lack reliable, training-free detection mechanisms. The authors propose a training-free spectral guardrail that identifies hallucinatory behavior by analyzing attention topology—specifically smoothness and entropy. Their findings reveal that spectral features from a single transformer layer can detect hallucinations with near-perfect accuracy, suggesting that hallucination fundamentally arises from abrupt shifts in the thermodynamic state of the model’s attention. The study also uncovers a “loud lies” phenomenon, wherein hallucinated outputs exhibit distinct spectral signatures. Evaluated on Llama 3.1 8B, the method achieves a 97.7% recall (98.2% using Layer 26 smoothness) and attains the best discriminative performance on Mistral 7B with an AUC of 0.900, demonstrating effective, label-free, and cross-model hallucination detection.

0 citationsRead paper
Recent publications

Latest Papers

A Probe Direction Is a Property of Its Prompt

Aug 13, 2026

This study addresses a critical oversight in existing probing methods for assessing whether language models are aware of being evaluated: the decisive influence of prompt selection on measurement outcomes, which undermines cross-model comparability. Treating prompts as a core component of measurement design, the authors fix task content while systematically varying prompts and employ controlled experiments, probe direction analysis, variance decomposition, and surface-form ablation tests to quantify the contributions of prompts, models, and their interaction to observed scores. Findings reveal that models account for only a small fraction of score variance; prompt choice can reverse apparent scaling trends; and surface-level prompt features alone suffice to reproduce most published results. The work demonstrates that single-prompt probing yields unreliable comparisons and specifies the minimum number of prompts required for robust evaluation.

0 citationsRead paper

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

Aug 13, 2026

This study addresses the lack of a standardized measurement location for latent feature importance in sparse autoencoder (SAE) evaluation, which distorts comparisons across dictionaries. Through controlled experiments, we demonstrate that variation in measurement location accounts for up to 11.9% of evaluation variance—a discrepancy that intensifies with larger corpora. To resolve this, we propose a causal evaluation framework based on ablation, which aligns latent features via decoder similarity, and introduce a simple yet effective standardization protocol: fixing the measurement location with a single line of code. This protocol disentangles dictionary performance from positional bias, and its necessity and efficacy are validated through multiple controlled experiments and audits of published works, substantially improving the consistency and comparability of SAE evaluations.

0 citationsRead paper

Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology

Feb 08, 2026

This work addresses the challenge of hallucination in autonomous agents deployed in open environments when using external tools, where existing approaches lack reliable, training-free detection mechanisms. The authors propose a training-free spectral guardrail that identifies hallucinatory behavior by analyzing attention topology—specifically smoothness and entropy. Their findings reveal that spectral features from a single transformer layer can detect hallucinations with near-perfect accuracy, suggesting that hallucination fundamentally arises from abrupt shifts in the thermodynamic state of the model’s attention. The study also uncovers a “loud lies” phenomenon, wherein hallucinated outputs exhibit distinct spectral signatures. Evaluated on Llama 3.1 8B, the method achieves a 97.7% recall (98.2% using Layer 26 smoothness) and attains the best discriminative performance on Mistral 7B with an AUC of 0.900, demonstrating effective, label-free, and cross-model hallucination detection.

0 citationsRead paper

Spectral Archaeology: The Causal Topology of Model Evolution

Jan 06, 2026arXiv.org

This work addresses the limitation of existing behavioral benchmarks, which capture only model outputs and fail to reveal internal mechanisms—particularly the structural discontinuities induced by curriculum shifts. The authors propose a training-free mechanistic probe that leverages spectral graph theory to analyze the algebraic connectivity (λ₂), smoothness, and spectral entropy of attention graphs, thereby constructing a “spectral fingerprint” that characterizes the causal topological structure underlying model evolution. They identify and name a novel phenomenon—“Passive-Triggered Connectivity Collapse” (PTCC)—demonstrating how curriculum changes disrupt syntactic sensitivity. A topology-based auditing framework is established and validated across 12 models and 10 languages, confirming fingerprint stability. Furthermore, sparse compensatory heads are localized to pinpoint failure mechanisms; guided activation of these heads restores approximately 38% of lost information flow, and the study reveals that topological structure is primarily governed by subword token density rather than language family.

0 citationsRead paper