Institution profile

Birmingham City University

Academic institutioneurope · gb
Official website
Research library19linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards Inclusive External Human-Machine Interface: Exploring the Effects of Visual and Auditory eHMI for Deaf and Hard-of-Hearing People

Jan 20, 2026

This study addresses the lack of inclusive design in external human-machine interfaces (eHMIs) for autonomous vehicles, which commonly overlook the communication needs of deaf and hard-of-hearing (DHH) individuals. It presents the first systematic investigation into eHMI usability for DHH users, employing focus group interviews, virtual reality simulations, eye-tracking, and subjective evaluations to compare visual and auditory eHMI modalities. Findings reveal that visual eHMIs significantly reduce crossing decision time and gaze duration for DHH participants while enhancing their trust, perceived safety, and perceived system usefulness, whereas auditory eHMIs yield limited effectiveness. Building on these insights, the study proposes five inclusive eHMI design principles tailored to DHH users, thereby addressing a critical gap in accessible intelligent transportation interaction research.

1 citationsRead paper

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Aug 14, 2026

This study addresses the insufficient robustness of temporal vision-language models against artifacts in endoscopic videos by proposing RobustEndoCLIP, a lightweight adaptation framework. Leveraging VeRA parameter-efficient fine-tuning and multimodal alignment techniques, this method enhances model resilience to visual disturbances. Furthermore, we introduce Endo-C6, the first corruption benchmark specifically designed for endoscopy, to enable systematic performance evaluation. Experimental results demonstrate that RobustEndoCLIP significantly improves both robustness and alignment accuracy under worst-case clinical artifact scenarios, comprehensively outperforming existing baselines. Consequently, this work establishes a reliable few-shot adaptation paradigm for intelligent endoscopic analysis, effectively bridging the gap between general-purpose vision-language models and specialized clinical applications requiring high reliability under challenging imaging conditions.

0 citationsRead paper

Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction

Jul 21, 2026

This study addresses the long-standing reliance on manual annotation for information extraction from architecture, engineering, and construction (AEC) drawings and the lack of effective research on layout detection in such domain-specific documents. To bridge this gap, the authors introduce the first AEC-focused drawing layout dataset and conduct a systematic evaluation of various deep learning models. Their analysis reveals, for the first time, a “domain interference” issue wherein general-purpose document understanding models underperform on AEC drawings. To overcome this limitation, they propose a novel approach combining the RF-DETR object detection model with the multimodal vision-language model Qwen3-VL. Experimental results demonstrate that RF-DETR achieves a mAP50 of 0.949 in layout detection, while Qwen3-VL attains an F1-score of 0.911 in information extraction, significantly outperforming existing general-purpose models and establishing both as state-of-the-art architectures for AEC drawing comprehension.

0 citationsRead paper
Recent publications

Latest Papers

On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

Aug 14, 2026

This study addresses the insufficient robustness of temporal vision-language models against artifacts in endoscopic videos by proposing RobustEndoCLIP, a lightweight adaptation framework. Leveraging VeRA parameter-efficient fine-tuning and multimodal alignment techniques, this method enhances model resilience to visual disturbances. Furthermore, we introduce Endo-C6, the first corruption benchmark specifically designed for endoscopy, to enable systematic performance evaluation. Experimental results demonstrate that RobustEndoCLIP significantly improves both robustness and alignment accuracy under worst-case clinical artifact scenarios, comprehensively outperforming existing baselines. Consequently, this work establishes a reliable few-shot adaptation paradigm for intelligent endoscopic analysis, effectively bridging the gap between general-purpose vision-language models and specialized clinical applications requiring high reliability under challenging imaging conditions.

0 citationsRead paper

Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction

Jul 21, 2026

This study addresses the long-standing reliance on manual annotation for information extraction from architecture, engineering, and construction (AEC) drawings and the lack of effective research on layout detection in such domain-specific documents. To bridge this gap, the authors introduce the first AEC-focused drawing layout dataset and conduct a systematic evaluation of various deep learning models. Their analysis reveals, for the first time, a “domain interference” issue wherein general-purpose document understanding models underperform on AEC drawings. To overcome this limitation, they propose a novel approach combining the RF-DETR object detection model with the multimodal vision-language model Qwen3-VL. Experimental results demonstrate that RF-DETR achieves a mAP50 of 0.949 in layout detection, while Qwen3-VL attains an F1-score of 0.911 in information extraction, significantly outperforming existing general-purpose models and establishing both as state-of-the-art architectures for AEC drawing comprehension.

0 citationsRead paper

Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation

Jul 13, 2026

This study addresses the challenges of unreliable probability calibration, inadequate uncertainty estimation, and the absence of a clinically actionable triage mechanism for thyroid nodule ultrasound images under distribution shift. To this end, the authors propose a deterministic classification framework based on a five-member deep ensemble, integrating a ConvNeXt-Tiny backbone, Squeeze-and-Excitation attention, member-wise vector scaling calibration, and mutual information–driven selective prediction to enable region-of-interest–level triage into three pathways: biopsy-free, biopsy-recommended, and radiologist review. On internal testing, the model achieves an AUC of 0.9395 (ECE = 0.0088), with a negative predictive value of 98.3% and a malignancy capture rate of 99.83% at a 50% retention rate in the biopsy-free pathway. External validation reveals a performance drop to an AUC of 0.7870, highlighting limitations in calibration threshold transferability while underscoring the method’s potential and challenges in high-stakes clinical settings.

0 citationsRead paper