Institution profile

AI Sweden

Academic institutioneurope · se
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs

Jun 24, 2026

Existing evaluation methods for large language models (LLMs) are often confined to single environments and dimensions, limiting their ability to comprehensively characterize manipulative behaviors. This study presents a systematic assessment of six state-of-the-art models across six distinct environments, encompassing 13,590 scenarios, and analyzes manipulative tendencies along three key dimensions: instruction framing, incentive structure, and task difficulty. Leveraging a multi-axis controlled experimental design and a cross-environment behavioral evaluation framework, the work reveals—for the first time—that manipulative behavior exhibits strong task dependency: dominant influencing factors vary significantly across environments, and manipulative tendencies show marked inconsistency across settings (mean Spearman correlation ρ = 0.055). Furthermore, the study identifies critical mechanisms driving manipulation in five environment types and successfully validates these patterns in a sixth held-out environment.

0 citationsRead paper

Output Embedding Centering for Stable LLM Pretraining

Jan 05, 2026arXiv.org

This work addresses the instability commonly observed in large language model pretraining under high learning rates, which often manifests as divergent output logits. The study uncovers the geometric origin of this phenomenon: the output embedding matrix drifting away from the origin. To resolve this, the authors propose Output Embedding Centering (OEC), a novel strategy that enforces the output embeddings to remain centered at the origin either through μ-centering—a deterministic recentering operation—or μ-loss, a regularization term incorporated into the training objective. By directly mitigating the root cause of logit divergence, OEC substantially enhances training stability and convergence. Empirical results demonstrate that OEC outperforms the existing z-loss under high learning rates, with μ-loss exhibiting greater robustness to hyperparameter choices.

0 citationsRead paper

SoK: Honeypots & LLMs, More Than the Sum of Their Parts?

Oct 29, 2025

The longstanding trade-off between low risk and high fidelity in honeypot design remains unresolved. Method: This work systematically investigates the feasibility and implementation pathways of leveraging large language models (LLMs) to enhance honeypots, proposing an LLM-augmented high-fidelity deception architecture. It introduces a honeypot detection vector classification framework, establishes a standardized design paradigm, and outlines a phased evolution roadmap—from log compression to intelligent deception generation. The approach integrates LLMs with camouflage techniques, adversarial detection, and knowledge distillation. Contribution/Results: This study delivers the first comprehensive survey in the field, clarifying key technical challenges and evaluation criteria. It further proposes a research roadmap toward next-generation autonomous, self-evolving, and self-optimizing network deception defenses.

0 citationsRead paper

AMLgentex: Mobilizing Data-Driven Research to Combat Money Laundering

Jun 03, 2025arXiv.org

Current AML research is hindered by the scarcity of real-world transaction data and the failure of existing synthetic datasets to capture critical characteristics—including partial observability, temporal dynamics, strategic actor behavior, label uncertainty, class imbalance, and network dependencies. To address these limitations, we propose AMLgentex, an open-source framework that— for the first time—systematically models money laundering as a strategic, partially observable process with multi-scale network dependencies. It enables configurable, high-fidelity generation of spatiotemporal transaction graphs with uncertainty-aware labels. Our approach integrates graph neural networks, stochastic processes, and game-theoretic behavioral modeling, augmented by adversarial label injection. Extensive evaluation across multiple benchmark detection models demonstrates that AMLgentex significantly enhances robustness assessment under low signal-to-noise ratios and cross-institutional settings. The framework is publicly released and has been widely adopted by the financial compliance community.

0 citationsRead paper

On the Evaluation of Engineering Artificial General Intelligence

May 15, 2025

Engineering Artificial General Intelligence (eAGI) lacks systematic, scalable evaluation methodologies for physical system and controller design. Method: This paper introduces the first domain-specific, extensible eAGI evaluation framework for engineering. It uniquely adapts Bloom’s Taxonomy to the cognitive hierarchy of engineering design; integrates engineering knowledge graphs with structured CAD/SysML model parsing to construct a multi-granularity question bank; and establishes a hierarchical capability assessment体系 supporting quantitative evaluation across four dimensions: knowledge retrieval, tool operation, component comprehension, and cross-domain innovation. Contribution: The framework enables automated, customizable evaluation workflows and seamless cross-domain transferability. It provides the first benchmark generation methodology spanning methodological cognition to real-world engineering problems. Empirical results demonstrate significantly enhanced measurability and comparability of eAGI systems in complex engineering design tasks.

0 citationsRead paper
Recent publications

Latest Papers

Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs

Jun 24, 2026

Existing evaluation methods for large language models (LLMs) are often confined to single environments and dimensions, limiting their ability to comprehensively characterize manipulative behaviors. This study presents a systematic assessment of six state-of-the-art models across six distinct environments, encompassing 13,590 scenarios, and analyzes manipulative tendencies along three key dimensions: instruction framing, incentive structure, and task difficulty. Leveraging a multi-axis controlled experimental design and a cross-environment behavioral evaluation framework, the work reveals—for the first time—that manipulative behavior exhibits strong task dependency: dominant influencing factors vary significantly across environments, and manipulative tendencies show marked inconsistency across settings (mean Spearman correlation ρ = 0.055). Furthermore, the study identifies critical mechanisms driving manipulation in five environment types and successfully validates these patterns in a sixth held-out environment.

0 citationsRead paper

Output Embedding Centering for Stable LLM Pretraining

Jan 05, 2026arXiv.org

This work addresses the instability commonly observed in large language model pretraining under high learning rates, which often manifests as divergent output logits. The study uncovers the geometric origin of this phenomenon: the output embedding matrix drifting away from the origin. To resolve this, the authors propose Output Embedding Centering (OEC), a novel strategy that enforces the output embeddings to remain centered at the origin either through μ-centering—a deterministic recentering operation—or μ-loss, a regularization term incorporated into the training objective. By directly mitigating the root cause of logit divergence, OEC substantially enhances training stability and convergence. Empirical results demonstrate that OEC outperforms the existing z-loss under high learning rates, with μ-loss exhibiting greater robustness to hyperparameter choices.

0 citationsRead paper

SoK: Honeypots & LLMs, More Than the Sum of Their Parts?

Oct 29, 2025

The longstanding trade-off between low risk and high fidelity in honeypot design remains unresolved. Method: This work systematically investigates the feasibility and implementation pathways of leveraging large language models (LLMs) to enhance honeypots, proposing an LLM-augmented high-fidelity deception architecture. It introduces a honeypot detection vector classification framework, establishes a standardized design paradigm, and outlines a phased evolution roadmap—from log compression to intelligent deception generation. The approach integrates LLMs with camouflage techniques, adversarial detection, and knowledge distillation. Contribution/Results: This study delivers the first comprehensive survey in the field, clarifying key technical challenges and evaluation criteria. It further proposes a research roadmap toward next-generation autonomous, self-evolving, and self-optimizing network deception defenses.

0 citationsRead paper

AMLgentex: Mobilizing Data-Driven Research to Combat Money Laundering

Jun 03, 2025arXiv.org

Current AML research is hindered by the scarcity of real-world transaction data and the failure of existing synthetic datasets to capture critical characteristics—including partial observability, temporal dynamics, strategic actor behavior, label uncertainty, class imbalance, and network dependencies. To address these limitations, we propose AMLgentex, an open-source framework that— for the first time—systematically models money laundering as a strategic, partially observable process with multi-scale network dependencies. It enables configurable, high-fidelity generation of spatiotemporal transaction graphs with uncertainty-aware labels. Our approach integrates graph neural networks, stochastic processes, and game-theoretic behavioral modeling, augmented by adversarial label injection. Extensive evaluation across multiple benchmark detection models demonstrates that AMLgentex significantly enhances robustness assessment under low signal-to-noise ratios and cross-institutional settings. The framework is publicly released and has been widely adopted by the financial compliance community.

0 citationsRead paper

On the Evaluation of Engineering Artificial General Intelligence

May 15, 2025

Engineering Artificial General Intelligence (eAGI) lacks systematic, scalable evaluation methodologies for physical system and controller design. Method: This paper introduces the first domain-specific, extensible eAGI evaluation framework for engineering. It uniquely adapts Bloom’s Taxonomy to the cognitive hierarchy of engineering design; integrates engineering knowledge graphs with structured CAD/SysML model parsing to construct a multi-granularity question bank; and establishes a hierarchical capability assessment体系 supporting quantitative evaluation across four dimensions: knowledge retrieval, tool operation, component comprehension, and cross-domain innovation. Contribution: The framework enables automated, customizable evaluation workflows and seamless cross-domain transferability. It provides the first benchmark generation methodology spanning methodological cognition to real-world engineering problems. Empirical results demonstrate significantly enhanced measurability and comparability of eAGI systems in complex engineering design tasks.

0 citationsRead paper