Institution profile

Edge Hill University

Academic institutioneurope · gb
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

Aug 13, 2026

This work addresses the challenge of fine-grained human action recognition, where visual similarity among actions impedes discriminability. To this end, we propose a structured multimodal fusion framework that jointly models three complementary representations: RGB appearance, pose heatmap geometry, and skeletal graph topology. Our approach employs pairwise cross-attention to enable symmetric inter-stream interaction and introduces a stream-wise latent sparse Mixture-of-Experts (MoE) mechanism that dynamically routes inputs to a shared subset of experts based on content, augmented with load-balancing regularization. Notably, our method achieves state-of-the-art performance on Gym99, Gym288, and Diving48 without relying on textual supervision or large-scale vision-language pretraining. On the long-tailed Gym288 benchmark, it improves mean class accuracy by 7.6 percentage points, from 68.6% to 76.2%.

0 citationsRead paper

Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

Aug 11, 2026

This study addresses the challenge of learning threat-aware and adaptive driving policies in safety-critical scenarios using only observable kinematic information and natural language descriptions. The authors propose a language-structured relational Q-learning approach, introducing language-guided scene modeling into relational reinforcement learning for the first time. They design an ego-centric relational Q-network (ERQ-Net) to jointly learn inter-vehicle dynamic dependencies and action values. The work formally identifies and articulates the “perception-control gap” problem, demonstrating improved performance in 2,500 CARLA simulation scenarios—raising success rates from 49–52% to 55–58% and enhancing adversarial target attention by 1.2–2.1×. Nevertheless, experiments reveal that 76% of scenarios remain solvable by simple policy compositions, highlighting a fundamental performance bottleneck in current methods.

0 citationsRead paper

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

Jun 18, 2026

This work addresses the challenges posed by high structural heterogeneity, large intra-class variation, and subtle visual differences between benign and malignant lesions in dermoscopic images. To this end, the authors propose a superpixel-based multimodal fusion approach that models lesions as graphs whose nodes correspond to superpixels. Node features are extracted using a frozen CNN, while geometric relationships are incorporated as edge attributes. A novel metadata context node is introduced to enable native graph-level fusion of clinical information with visual features. Discriminative classification embeddings are generated through an edge-aware Graph Transformer coupled with an attention propagation mechanism. This method, which uniquely integrates superpixel graph structure, geometric edge attributes, and metadata context, achieves state-of-the-art performance across four public datasets, significantly improving both accuracy and robustness in benign–malignant skin lesion classification.

0 citationsRead paper
Recent publications

Latest Papers

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

Aug 13, 2026

This work addresses the challenge of fine-grained human action recognition, where visual similarity among actions impedes discriminability. To this end, we propose a structured multimodal fusion framework that jointly models three complementary representations: RGB appearance, pose heatmap geometry, and skeletal graph topology. Our approach employs pairwise cross-attention to enable symmetric inter-stream interaction and introduces a stream-wise latent sparse Mixture-of-Experts (MoE) mechanism that dynamically routes inputs to a shared subset of experts based on content, augmented with load-balancing regularization. Notably, our method achieves state-of-the-art performance on Gym99, Gym288, and Diving48 without relying on textual supervision or large-scale vision-language pretraining. On the long-tailed Gym288 benchmark, it improves mean class accuracy by 7.6 percentage points, from 68.6% to 76.2%.

0 citationsRead paper

Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

Aug 11, 2026

This study addresses the challenge of learning threat-aware and adaptive driving policies in safety-critical scenarios using only observable kinematic information and natural language descriptions. The authors propose a language-structured relational Q-learning approach, introducing language-guided scene modeling into relational reinforcement learning for the first time. They design an ego-centric relational Q-network (ERQ-Net) to jointly learn inter-vehicle dynamic dependencies and action values. The work formally identifies and articulates the “perception-control gap” problem, demonstrating improved performance in 2,500 CARLA simulation scenarios—raising success rates from 49–52% to 55–58% and enhancing adversarial target attention by 1.2–2.1×. Nevertheless, experiments reveal that 76% of scenarios remain solvable by simple policy compositions, highlighting a fundamental performance bottleneck in current methods.

0 citationsRead paper

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

Jun 18, 2026

This work addresses the challenges posed by high structural heterogeneity, large intra-class variation, and subtle visual differences between benign and malignant lesions in dermoscopic images. To this end, the authors propose a superpixel-based multimodal fusion approach that models lesions as graphs whose nodes correspond to superpixels. Node features are extracted using a frozen CNN, while geometric relationships are incorporated as edge attributes. A novel metadata context node is introduced to enable native graph-level fusion of clinical information with visual features. Discriminative classification embeddings are generated through an edge-aware Graph Transformer coupled with an attention propagation mechanism. This method, which uniquely integrates superpixel graph structure, geometric edge attributes, and metadata context, achieves state-of-the-art performance across four public datasets, significantly improving both accuracy and robustness in benign–malignant skin lesion classification.

0 citationsRead paper