Institution profile

Tübingen AI Center

Academic institutioneurope · de
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

Jun 24, 2026

This work investigates how to reverse-engineer executable decision code of an agent solely from its behavioral trajectories in games and examines how actively designed adversarial experiments can enhance reconstruction fidelity. To this end, we introduce RevengeBench, a benchmark comprising 75 Elo-calibrated CodeClash strategies across five environments, where customized behavioral probes are generated by observing interactions between target and opponent agents to reconstruct underlying code logic. We formalize strategy reversal as a tractable inverse problem in code space and incorporate a mechanism for controlled experimentation. Experiments across 12 large language models demonstrate that our approach significantly reduces initial behavioral divergence (by 34%–72%) and yields reconstructed strategies that exhibit competitive performance in downstream adversarial settings, particularly bolstering the counterplay capabilities of weaker models.

0 citationsRead paper

Un-Attributability: Computing Novelty From Retrieval & Semantic Similarity

Oct 31, 2025

This work addresses the challenge of quantifying semantic novelty in language model outputs, proposing *unattributability*—the inability to semantically retrieve any pretraining corpus sample as the source of an output—as a formal, interpretable metric. Methodologically, it introduces a two-stage retrieval pipeline: first, efficient coarse-grained indexing using GIST embeddings; second, fine-grained re-ranking via ColBERTv2, with attribution thresholds calibrated against human-written text. The paper provides the first formal definition and empirical evaluation of unattributability. Experiments on the SmolLM family reveal three key findings: (1) instruction tuning significantly improves output unattributability; (2) increased reliance on longer contexts enhances semantic novelty; and (3) domain-specific characteristics shape the distribution of unattributable outputs. By grounding novelty assessment in semantic retrieval fidelity rather than surface-level heuristics, this work establishes a scalable, principled framework for evaluating generative originality—offering both interpretability and practical applicability for safety, copyright, and alignment research.

0 citationsRead paper
Recent publications

Latest Papers

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

Jun 24, 2026

This work investigates how to reverse-engineer executable decision code of an agent solely from its behavioral trajectories in games and examines how actively designed adversarial experiments can enhance reconstruction fidelity. To this end, we introduce RevengeBench, a benchmark comprising 75 Elo-calibrated CodeClash strategies across five environments, where customized behavioral probes are generated by observing interactions between target and opponent agents to reconstruct underlying code logic. We formalize strategy reversal as a tractable inverse problem in code space and incorporate a mechanism for controlled experimentation. Experiments across 12 large language models demonstrate that our approach significantly reduces initial behavioral divergence (by 34%–72%) and yields reconstructed strategies that exhibit competitive performance in downstream adversarial settings, particularly bolstering the counterplay capabilities of weaker models.

0 citationsRead paper

Un-Attributability: Computing Novelty From Retrieval & Semantic Similarity

Oct 31, 2025

This work addresses the challenge of quantifying semantic novelty in language model outputs, proposing *unattributability*—the inability to semantically retrieve any pretraining corpus sample as the source of an output—as a formal, interpretable metric. Methodologically, it introduces a two-stage retrieval pipeline: first, efficient coarse-grained indexing using GIST embeddings; second, fine-grained re-ranking via ColBERTv2, with attribution thresholds calibrated against human-written text. The paper provides the first formal definition and empirical evaluation of unattributability. Experiments on the SmolLM family reveal three key findings: (1) instruction tuning significantly improves output unattributability; (2) increased reliance on longer contexts enhances semantic novelty; and (3) domain-specific characteristics shape the distribution of unattributable outputs. By grounding novelty assessment in semantic retrieval fidelity rather than surface-level heuristics, this work establishes a scalable, principled framework for evaluating generative originality—offering both interpretability and practical applicability for safety, copyright, and alignment research.

0 citationsRead paper