Institution profile

University of Manchester

Academic institutioneurope · gb
Official website
Research library853linked papers
Opportunities0open roles
Selected work

Representative Papers

OmniBench: Towards The Future of Universal Omni-Language Models

Sep 23, 2024arXiv.org

Existing open-source multimodal large language models (MLLMs) exhibit significant deficiencies in joint visual-auditory-textual understanding and reasoning, achieving only ~50% instruction-following accuracy on trilingual multimodal tasks. Method: We introduce OmniBench—the first benchmark for trilingual multimodal collaborative reasoning—and formalize the omni-language model (OLM), a unified architecture capable of jointly processing visual, auditory, and textual (V-A-T) inputs. We construct OmniBench via expert human annotation across diverse trilingual multimodal tasks and curate OmniInstruct, a large-scale instruction-tuning dataset comprising 96K samples. Our methodology integrates cross-modal alignment modeling, trilingual multimodal instruction tuning, and a human-in-the-loop evaluation framework. Contribution/Results: Experiments reveal severe generalization limitations of current open-source OLMs on trilingual multimodal tasks; OmniInstruct substantially improves their reasoning performance. This work establishes a novel evaluation paradigm, provides high-quality resources, and outlines a technical pathway for advancing trilingual multimodal foundation models.

9 citations2 influentialRead paper

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Jan 20, 2026

This work proposes a practical, three-stage “Locate–Guide–Improve” framework that transforms mechanistic interpretability from a post-hoc diagnostic tool into an engineering-driven optimization methodology for large language models. By systematically integrating techniques for identifying critical neurons and pathways with targeted interventions—such as activation manipulation and module editing—the framework establishes a standardized protocol for model refinement while clearly distinguishing between localization and guidance mechanisms. Empirical results demonstrate significant improvements in model alignment, task performance, and reasoning efficiency, thereby advancing mechanistic interpretability toward real-world applicability.

4 citationsRead paper

Causal Explanations for Image Classifiers

Nov 13, 2024arXiv.org

Existing explanation methods for image classifiers lack rigorous formal definitions of causality and explanation, relying predominantly on heuristic strategies. Method: This paper introduces the Halpern–Pearl theory of actual causality to black-box image classification interpretability—the first systematic application of this causal framework to the domain. We propose REX, a causally grounded explanation generation framework that formally defines “cause” and “explanation,” designs a provably terminating algorithm for approximating minimal explanations, and implements an iterative solving mechanism with controllable computational complexity. Contribution/Results: The implemented tool REX outperforms state-of-the-art black-box explanation methods across explanation compactness, computational efficiency, and standard quality metrics (e.g., fidelity, stability, and comprehensibility). Experiments demonstrate that REX produces the most concise explanations and achieves the fastest convergence. This work establishes a rigorous causal foundation for explainable AI while delivering a practical, scalable technical solution.

4 citationsRead paper

Natural Language Generation

Oct 24, 2018Theoretical Issues In Natural Language Processing

Natural language generation (NLG) lacks a unified conceptual framework and clearly delineated disciplinary boundaries, leading to ambiguity in its scope relative to other NLP subfields such as machine translation and dialogue systems. Method: This paper systematically surveys NLG’s research landscape and historical evolution, focusing on core tasks—including data-to-text generation, text summarization, and image captioning—and proposes a novel taxonomy grounded in the dual primitives of “content selection” and “realization.” It traces methodological shifts from rule-based and statistical approaches to early neural models, analyzing how task-driven evaluation paradigms have evolved. Contribution/Results: The work establishes NLG’s formal disciplinary boundaries for the first time, constructs a widely cited conceptual framework and discipline map, clarifies terminological consensus, and identifies large language models as catalysts for methodological convergence across NLP subfields. These contributions lay foundational theoretical groundwork for NLG research and practice.

2 citations1 influentialRead paper

PAMAS: Self-Adaptive Multi-Agent System with Perspective Aggregation for Misinformation Detection

Feb 03, 2026

This work addresses the challenges posed by the high diversity and context dependence of misinformation on social media, which often overwhelms conventional multi-agent systems and obscures subtle deceptive cues. To overcome these limitations, the authors propose an adaptive multi-agent framework featuring a tripartite role structure—auditor, coordinator, and decision-maker—augmented with a perspective-aware hierarchical aggregation mechanism and an adaptive topology optimization strategy. This design effectively amplifies anomalous signals while integrating heterogeneous viewpoints. Leveraging large language model–driven collaborative reasoning, dynamic routing, and an evolving memory mechanism, the approach significantly outperforms state-of-the-art methods across multiple benchmark datasets, achieving superior detection accuracy and reasoning efficiency. The proposed solution offers a scalable and robust paradigm for misinformation detection in complex social media environments.

1 citationsRead paper
Recent publications

Latest Papers