Institution profile

University of Lincoln

Academic institutioneurope · gb
Official website
Research library57linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards counterfactual and contrastive explainability and transparency of DCNN image classifiers

Sep 01, 2022Knowledge-Based Systems

To address the insufficient decision transparency of deep convolutional neural networks (DCNNs) in image classification, this paper proposes the first unified interpretability framework that jointly models counterfactual perturbations and class-contrastive mechanisms, generating human-understandable “why-this-not-that” explanations. Methodologically, it integrates gradient-guided counterfactual search, contrastive attention mask optimization, a differentiable semantic editing module, and a class-aware loss function—ensuring both local faithfulness and global consistency while enabling fine-grained attribution and controllable semantic editing. Experiments on ImageNet and CUB-200 demonstrate significant improvements: +23.6% in explanation fidelity and +31.2% in user trustworthiness. The generated explanations exhibit strong semantic plausibility and visual verifiability.

7 citationsRead paper

Counting with Confidence: Accurate Pest Monitoring in Water Traps

May 19, 2025IFAC-PapersOnLine

To address the challenge of unreliable pest counting in real-world deployment due to the absence of ground-truth annotations and difficulty in assessing result credibility, this paper proposes the first end-to-end confidence estimation framework for pest image counting. The method integrates multi-source features—including object detection outputs, image sharpness (measured by mean gradient magnitude), image quality and complexity, and spatial uniformity of pest distribution (quantified via adaptive DBSCAN)—into a multi-feature regression model. Crucially, it introduces a hypothesis-driven, multi-factor sensitivity analysis to identify the most discriminative evaluation metrics. Evaluated on a custom-built test set, the framework reduces mean squared error by 31.7% and improves the coefficient of determination (R²) by 15.2% over a baseline relying solely on detection outputs, demonstrating substantial gains in both accuracy and robustness of confidence estimation.

1 citationsRead paper

Mapping the Unseen: Unified Promptable Panoptic Mapping with Dynamic Labeling using Foundation Models

May 03, 2024arXiv.org

Traditional panoramic semantic mapping is constrained by predefined categories and struggles to recognize unknown objects. To address this, we propose the first promptable unified panoramic mapping framework. Our method integrates natural language prompts with multimodal foundation models (CLIP and SAM) to construct a prompt-driven dynamic labeling module, enabling real-time open-vocabulary semantic parsing. By jointly leveraging 3D reconstruction and instance segmentation, the framework achieves end-to-end promptable panoramic mapping. Extensive evaluation on both real-world and synthetic datasets demonstrates significant improvements in unknown-object segmentation accuracy and semantic labeling fidelity, while enabling natural-language-guided interactive map construction. Ablation studies further validate that foundation-model-based dynamic labeling substantially outperforms conventional fixed-label paradigms. This work establishes a new paradigm for flexible, scalable, and user-controllable semantic mapping beyond closed-set assumptions.

1 citationsRead paper

Detecting and Deterring Manipulation in a Cognitive Hierarchy

May 03, 2024

Intelligent agents with limited nested reasoning—such as low-order agents in Interactive Partially Observable Markov Decision Processes (IPOMDPs)—are vulnerable to manipulation by higher-order adversaries; existing recursive modeling frameworks struggle to simultaneously ensure interpretability and robust countermeasures. Method: We propose the ℵ-IPOMDP framework, the first to integrate statistical anomaly detection with *out-of-belief* policies within the IPOMDP formalism. This enables low-order agents to detect deceptive behavior and enact credible deterrence without requiring explicit understanding of higher-order reasoning mechanisms. Contribution/Results: ℵ-IPOMDP significantly reduces the success rate of higher-order exploitation in both mixed-motive and zero-sum games, thereby enhancing interaction fairness. It provides a lightweight, deployable robust adversarial mechanism for AI safety, cybersecurity, and cognitive modeling—balancing computational efficiency, interpretability, and resilience against strategic deception.

1 citationsRead paper

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

Aug 08, 2026

This work addresses whether embodied vision-language models (VLMs) possess reliable spatial understanding and decision-making capabilities in partially observable, safety-critical environments by proposing the Explore, Map, Remember, Decide (EMRD) evaluation framework. EMRD extends theoretical spatial cognition into safety-critical domains for the first time, assessing exploration through environmental coverage and temporal efficiency, while incorporating metrics for spatial mapping fidelity, psychological memory tests, and focal decision-making. Robustness is further evaluated under low-light conditions and texture perturbations. The study reveals that VLMs often rely on textual priors rather than spatial evidence when selecting evacuation points, exhibit significantly degraded spatial reasoning in low light yet remain robust to texture interference, and employ memory mechanisms fundamentally divergent from human cognition—introducing unpredictable alignment risks.

0 citationsRead paper
Recent publications

Latest Papers

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

Aug 08, 2026

This work addresses whether embodied vision-language models (VLMs) possess reliable spatial understanding and decision-making capabilities in partially observable, safety-critical environments by proposing the Explore, Map, Remember, Decide (EMRD) evaluation framework. EMRD extends theoretical spatial cognition into safety-critical domains for the first time, assessing exploration through environmental coverage and temporal efficiency, while incorporating metrics for spatial mapping fidelity, psychological memory tests, and focal decision-making. Robustness is further evaluated under low-light conditions and texture perturbations. The study reveals that VLMs often rely on textual priors rather than spatial evidence when selecting evacuation points, exhibit significantly degraded spatial reasoning in low light yet remain robust to texture interference, and employ memory mechanisms fundamentally divergent from human cognition—introducing unpredictable alignment risks.

0 citationsRead paper

Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement

Jul 30, 2026

This study addresses the critical issue that branch disconnections in coronary artery segmentation lead to erroneous FFR-CT–based treatment decisions, a flaw undetected by conventional metrics such as Dice coefficient due to their insensitivity to topological connectivity. To bridge this gap, the work introduces Bifurcation Connectedness Score (BCS)—the first dedicated metric for assessing bifurcation integrity—and its differentiable variant, soft-BCS, to explicitly evaluate and optimize the topological fidelity of segmentation outputs. Validation through centerline analysis, FFR-CT hemodynamic simulations, and deep learning training demonstrates that BCS effectively disentangles two distinct properties: “branch recovery” and “connection preservation.” Experimental results show that higher BCS significantly improves agreement between FFR-CT–derived clinical decisions based on predicted versus ground-truth geometries (OR = 2.16), with the most pronounced benefits observed in severe stenosis cases.

0 citationsRead paper

Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation

Jun 20, 2026

This work addresses the challenge of efficiently adapting large vision foundation models to cell segmentation, which typically requires computationally expensive fine-tuning and abundant annotated data. The authors propose EffiCell-Seg, a novel framework that operates with a frozen pretrained vision encoder and introduces, for the first time, an insight into its inherent complementary priors—global saliency and local morphology. Leveraging this observation, they design a training-free adaptation mechanism: a Cell Structure Prompt Encoder generates structural prior maps, which are synergistically integrated with a Mask Decoder that mutually guides predictions using geometric distance fields and semantic maps. Evaluated across diverse cellular imaging modalities, EffiCell-Seg achieves state-of-the-art performance while employing only approximately 5 million trainable parameters—over 130 times fewer than full fine-tuning.

0 citationsRead paper

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

May 26, 2026

This study addresses the scarcity of high-quality speech data that limits the training and evaluation of large language models (LLMs) in real-world medical spoken dialogues. To bridge this gap, the authors introduce MeDial-Speech, a novel dataset comprising over 111 hours of authentic clinician–patient and robot–patient conversational speech across four medical conditions, accompanied by both human and automatic transcripts with word error rates yielding accuracies of 71.1% and 74.7%, respectively. The work further proposes a sentence-selection-based paradigm for evaluating dialogue understanding and benchmarks prominent LLMs—including GPT-5 Mini, DeepSeek-V3, and Claude Sonnet 4—finding that Claude Sonnet 4 achieves the best performance, yet all models exhibit significant overconfidence in their predictions. This resource establishes a foundational benchmark for advancing spoken medical dialogue systems.

0 citationsRead paper

Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection

May 24, 2026

This work addresses the limitation of existing 3D bird’s-eye-view (BEV) object detection methods that employ uniform random masking during multimodal pretraining, which neglects semantically critical regions and hampers representation learning. To overcome this, the authors propose a Semantic-Guided Multimodal Masked Autoencoder (SG-M2AE) that integrates semantic priors into pretraining. Specifically, they design a semantic-guided LiDAR voxel masking strategy that preferentially preserves regions with high semantic value and introduce a point-level semantic decoder head as an auxiliary supervision signal to enhance cross-modal representation learning. Evaluated on the nuScenes mini validation set, SG-M2AE significantly outperforms the UniM2AE baseline, achieving a 1.49% improvement in mAP and a 3.22% gain in NDS, thereby demonstrating the effectiveness and novelty of the proposed semantic-guided mechanism.

0 citationsRead paper