Institution profile

Silo AI

Industry researcheurope · fi
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Jun 22, 2026

Existing multimodal large language models typically perform coarse-grained representation alignment only at fixed linguistic layers, overlooking the fine-grained structure within individual Transformer attention heads, which limits cross-modal alignment efficacy. This work proposes HeRA, the first method to achieve cross-modal topological alignment at the level of individual attention heads. Grounded in the Bregman representation hypothesis, HeRA enhances alignment quality by preserving local neighborhood relationships across modalities. It introduces a differentiable Mutual K-Nearest Neighbor (MKNN) contrastive objective that dynamically identifies and optimizes critical attention heads, revealing that the least-aligned heads yield the greatest performance gains. Experiments demonstrate that HeRA significantly improves visual-centric task performance across multiple state-of-the-art multimodal large language models and 18 benchmarks, while effectively mitigating visual hallucination and overreliance on linguistic priors.

0 citationsRead paper

Manifesto from Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE)

Jun 08, 2025

Current conversational information access (CONIAC) systems lack a unified, human-centered evaluation framework. Method: This paper introduces CAFE—the first consensus-driven, multidimensional evaluation framework for CONIAC—systematically integrating six core dimensions: stakeholder goals, user tasks, user characteristics, evaluation criteria, methodologies, and quantitative metrics. Moving beyond traditional technology-centric, unidimensional paradigms, CAFE employs world-model abstraction and interdisciplinary consensus workshops—drawing on human-computer interaction, information retrieval, and evaluation science—to structurally align system objectives with human factors. Contribution/Results: As the first internationally recognized, consensus-driven evaluation framework for conversational agents, CAFE has been formally published in the Dagstuhl Perspectives Workshop series. It provides both theoretical foundations and practical guidelines for the design, evaluation, and standardization of CONIAC systems.

0 citationsRead paper

DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

Apr 28, 2025

Conventional multimodal emotion recognition relies on identity-sensitive cues (e.g., facial appearance and voiceprint), posing significant privacy risks. Method: This paper proposes a de-identified multimodal emotion recognition and reasoning paradigm. We introduce the first dual-modal benchmark dataset supporting Non-Facial Body Language (NFBL) modeling and instruction-driven reasoning, and design DEEMO-LLaMA—a unified architecture integrating de-identified audio-visual representation learning, an NFBL-aware module, and a cue-alignment reasoning mechanism. Crucially, it achieves identity-agnostic emotion understanding without facial or vocal biometric cues. Contribution/Results: Our method attains 74.49% accuracy (F1 = 74.45%) on emotion classification; for reasoning tasks, it achieves cue–label overlap scores of 6.20 and 7.66—substantially outperforming existing Multimodal Large Language Models (MLLMs). This work establishes a foundational framework for privacy-preserving, trustworthy affective computing in sensitive applications.

0 citationsRead paper

SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes

Apr 16, 2025

Hallucinations and over-generation in multilingual large language models (LLMs) remain poorly understood and inadequately evaluated across languages. Method: This work introduces the first large-scale, shared multilingual hallucination detection task covering 14 languages, formalized as a cross-lingual span-level annotation problem. We propose a unified multilingual hallucination span annotation framework, uncovering cross-lingual disparities in hallucination distribution and annotator disagreement—thereby advancing detection from coarse-grained binary classification toward fine-grained localization. Contribution/Results: Leveraging span-level annotations, multilingual LLM output evaluation, and cross-lingual benchmarking, we organized a competition with 43 teams submitting 2,618 systems, establishing the first authoritative multilingual hallucination detection baseline. Empirical analysis identifies model scale, language resource coverage, and post-processing strategies as the three primary determinants of detection performance.

0 citationsRead paper

Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation

Mar 27, 2025

To address the severe cross-domain performance degradation in open-vocabulary semantic segmentation, this paper proposes SemLA—a training-free test-time domain adaptation framework. SemLA constructs a LoRA adapter semantic library and leverages CLIP’s semantic space alignment to perform nearest-neighbor retrieval and weighted fusion of adapters for each input image, dynamically synthesizing an image-level personalized segmentation model. It establishes the first “zero-training, semantics-driven” domain adaptation paradigm for open-vocabulary segmentation, offering inherent interpretability (via adapter contribution tracing), strong data privacy preservation, and efficient scalability. Evaluated across a comprehensive benchmark spanning 10 datasets and 20 domains, SemLA significantly outperforms existing methods, setting a new state-of-the-art for domain-adaptive open-vocabulary semantic segmentation.

0 citationsRead paper
Recent publications

Latest Papers

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

Jun 22, 2026

Existing multimodal large language models typically perform coarse-grained representation alignment only at fixed linguistic layers, overlooking the fine-grained structure within individual Transformer attention heads, which limits cross-modal alignment efficacy. This work proposes HeRA, the first method to achieve cross-modal topological alignment at the level of individual attention heads. Grounded in the Bregman representation hypothesis, HeRA enhances alignment quality by preserving local neighborhood relationships across modalities. It introduces a differentiable Mutual K-Nearest Neighbor (MKNN) contrastive objective that dynamically identifies and optimizes critical attention heads, revealing that the least-aligned heads yield the greatest performance gains. Experiments demonstrate that HeRA significantly improves visual-centric task performance across multiple state-of-the-art multimodal large language models and 18 benchmarks, while effectively mitigating visual hallucination and overreliance on linguistic priors.

0 citationsRead paper

Manifesto from Dagstuhl Perspectives Workshop 24352 -- Conversational Agents: A Framework for Evaluation (CAFE)

Jun 08, 2025

Current conversational information access (CONIAC) systems lack a unified, human-centered evaluation framework. Method: This paper introduces CAFE—the first consensus-driven, multidimensional evaluation framework for CONIAC—systematically integrating six core dimensions: stakeholder goals, user tasks, user characteristics, evaluation criteria, methodologies, and quantitative metrics. Moving beyond traditional technology-centric, unidimensional paradigms, CAFE employs world-model abstraction and interdisciplinary consensus workshops—drawing on human-computer interaction, information retrieval, and evaluation science—to structurally align system objectives with human factors. Contribution/Results: As the first internationally recognized, consensus-driven evaluation framework for conversational agents, CAFE has been formally published in the Dagstuhl Perspectives Workshop series. It provides both theoretical foundations and practical guidelines for the design, evaluation, and standardization of CONIAC systems.

0 citationsRead paper

DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

Apr 28, 2025

Conventional multimodal emotion recognition relies on identity-sensitive cues (e.g., facial appearance and voiceprint), posing significant privacy risks. Method: This paper proposes a de-identified multimodal emotion recognition and reasoning paradigm. We introduce the first dual-modal benchmark dataset supporting Non-Facial Body Language (NFBL) modeling and instruction-driven reasoning, and design DEEMO-LLaMA—a unified architecture integrating de-identified audio-visual representation learning, an NFBL-aware module, and a cue-alignment reasoning mechanism. Crucially, it achieves identity-agnostic emotion understanding without facial or vocal biometric cues. Contribution/Results: Our method attains 74.49% accuracy (F1 = 74.45%) on emotion classification; for reasoning tasks, it achieves cue–label overlap scores of 6.20 and 7.66—substantially outperforming existing Multimodal Large Language Models (MLLMs). This work establishes a foundational framework for privacy-preserving, trustworthy affective computing in sensitive applications.

0 citationsRead paper

SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes

Apr 16, 2025

Hallucinations and over-generation in multilingual large language models (LLMs) remain poorly understood and inadequately evaluated across languages. Method: This work introduces the first large-scale, shared multilingual hallucination detection task covering 14 languages, formalized as a cross-lingual span-level annotation problem. We propose a unified multilingual hallucination span annotation framework, uncovering cross-lingual disparities in hallucination distribution and annotator disagreement—thereby advancing detection from coarse-grained binary classification toward fine-grained localization. Contribution/Results: Leveraging span-level annotations, multilingual LLM output evaluation, and cross-lingual benchmarking, we organized a competition with 43 teams submitting 2,618 systems, establishing the first authoritative multilingual hallucination detection baseline. Empirical analysis identifies model scale, language resource coverage, and post-processing strategies as the three primary determinants of detection performance.

0 citationsRead paper

Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation

Mar 27, 2025

To address the severe cross-domain performance degradation in open-vocabulary semantic segmentation, this paper proposes SemLA—a training-free test-time domain adaptation framework. SemLA constructs a LoRA adapter semantic library and leverages CLIP’s semantic space alignment to perform nearest-neighbor retrieval and weighted fusion of adapters for each input image, dynamically synthesizing an image-level personalized segmentation model. It establishes the first “zero-training, semantics-driven” domain adaptation paradigm for open-vocabulary segmentation, offering inherent interpretability (via adapter contribution tracing), strong data privacy preservation, and efficient scalability. Evaluated across a comprehensive benchmark spanning 10 datasets and 20 domains, SemLA significantly outperforms existing methods, setting a new state-of-the-art for domain-adaptive open-vocabulary semantic segmentation.

0 citationsRead paper