Institution profile

University of Potsdam

Academic institutioneurope · de
Official website
Research library276linked papers
Opportunities0open roles
Selected work

Representative Papers

Manipulating Feature Visualizations with Gradient Slingshots

Jan 11, 2024arXiv.org

This work exposes a critical credibility vulnerability in feature visualization (FV) for deep neural network interpretability: FV outputs are susceptible to stealthy manipulation, leading to erroneous attribution of neuron semantics. To address this, we propose the first model-architecture-agnostic targeted FV manipulation method. Our approach integrates gradient redirection (via Slingshot optimization), adversarial latent-space perturbations, and neuron-activation-constrained regularization to achieve “semantic masking”—i.e., seamless substitution of a target neuron’s original FV explanation with an arbitrary user-specified semantic concept. Experiments across CNNs and Vision Transformers demonstrate successful concealment of functionally critical neurons: model accuracy degrades by less than 0.3%, yet FV-based auditing yields a 92% false-negative rate in detecting manipulated neurons. These results underscore the fragility of prevailing FV techniques and establish a new paradigm for robust model auditing and interpretability governance.

6 citationsRead paper

Causal inference for N-of-1 trials

Jun 14, 2024

This study addresses personalized causal inference in N-of-1 trials (within-subject crossover experiments) by proposing the first causal framework tailored to individual subjects. Methodologically, it (1) formally establishes identifiability conditions for causal effects in N-of-1 trials; (2) defines and estimates the unit-level conditional average treatment effect (U-CATE) to capture dynamic, time-varying individual causal mechanisms; and (3) develops a g-formula-based identification strategy for U-CATE under time-varying confounding and residual carryover effects, accompanied by theoretical guarantees. We prove that, under standard assumptions, the simple mean-difference estimator is consistent for U-CATE. Empirical analysis on acne N-of-1 trial data demonstrates substantial estimation discrepancies across modeling assumptions, underscoring the importance of appropriate causal identification. The framework significantly enhances the reliability and interpretability of individualized treatment decisions.

3 citations1 influentialRead paper

Towards a Theory on Process Automation Effects

Mar 25, 2025arXiv.org

Prior research predominantly focuses on the design and deployment of process automation, neglecting its real-world operational impacts after implementation. Method: This paper addresses this gap through a systematic literature review of human–machine collaboration, constructing the first theoretical framework specifically for *in-production* process automation. It proposes a novel four-part dynamic co-adaptation model—comprising technology, participants, managers, and developers—that transcends traditional binary (human/machine) analytical paradigms. Leveraging cross-domain theoretical integration and conceptual modeling, the study establishes a transferable framework for evaluating automation outcomes. Contribution/Results: The framework yields actionable pathways for organizational optimization of automation practices and identifies several novel research questions, thereby advancing a coherent, systemic research agenda for in-production automation in both academia and practice.

3 citationsRead paper

Step-resolved data attribution for looped transformers

Feb 10, 2026

This work addresses the challenge of characterizing the influence of individual training samples across the iterative steps of recurrent Transformers, a capability lacking in existing data influence estimation methods. To this end, the authors propose Step-Decomposed Influence (SDI), which unfolds the recurrent computation graph to decompose data influence at each inference step. By integrating the TracIn framework with TensorSketch approximation, SDI avoids explicit per-sample gradient computation, enabling efficient and scalable fine-grained attribution. Experiments demonstrate that SDI achieves high accuracy and strong scalability on recurrent GPT models and algorithmic reasoning tasks, facilitating multi-dimensional interpretability analyses of the internal reasoning dynamics within recurrent architectures.

1 citationsRead paper

Craig Interpolation for HT with a Variation of Mints' Sequent System

Jan 07, 2026arXiv.org

This work addresses the construction of Craig interpolants for three-valued logic of Here and There (HT), also known as Gödel logic G₃. To this end, it proposes a two-stage approach: first, an initial interpolant is constructed within a generalized non-classical logic enriched with auxiliary operators, and then it is transformed into a valid HT interpolant. The key innovation lies in the first adaptation of a Maehara-style interpolation method to HT logic, directly operating on HT formulas through a variant of Mints’ sequent calculus. This approach successfully yields Craig interpolants in HT logic, thereby demonstrating the feasibility and effectiveness of interpolation techniques in this non-classical setting.

1 citationsRead paper
Recent publications

Latest Papers