Institution profile

University of Beira Interior

Academic institutioneurope · pt
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device

Aug 10, 2026

This study addresses the inefficiency and low accuracy of manual length measurement for industrial cable reels by proposing an end-to-end vision-based method that estimates cable length from a single smartphone image. For the first time, camera calibration, instance segmentation, 6D pose estimation, and volume computation are integrated into a handheld device pipeline. The approach accommodates diverse reel types and cable diameters, enabling rapid length estimation from just one image. Trained on 1,000 annotated images, the instance segmentation model achieves a test-set mAP50 of 99.5%. Evaluated on 75 real-world cable reels, the system yields an average absolute percentage error of 4.90% with a per-image inference time of only 5.66 milliseconds, significantly enhancing both efficiency and precision in field measurements.

0 citationsRead paper

VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

Jul 07, 2026

This work addresses the lack of a unified evaluation protocol that hinders direct comparison among deepfake detection approaches across three dominant paradigms: commercial APIs, zero-shot vision-language models, and open-source detectors. To bridge this gap, the authors introduce VendorBench-100, a benchmark comprising a standardized corpus of 100 adversarial images, evaluated under consistent output formatting and a dual-metric strategy prioritizing the Matthews Correlation Coefficient (MCC) with ROC-AUC as a secondary measure. For the first time, 36 representative models are systematically assessed within a common framework. The results reveal that commercial APIs generally achieve the best performance, while certain open-source models rival top-tier vision-language models. Notably, high ROC-AUC scores do not necessarily correspond to high MCC values, underscoring the critical influence of metric selection on evaluation outcomes.

0 citationsRead paper

Robust Face Super-Resolution and Recognition Through Multi-Feature Aggregation in Diffusion Models

Jul 06, 2026

This work addresses the challenge of face super-resolution in surveillance scenarios, where low resolution, pose variations, uneven illumination, and occlusions severely hinder recognition performance. Conventional super-resolution methods often introduce identity distortions that degrade downstream tasks. To overcome this, the authors propose FASR++, a novel approach that, for the first time, integrates features from a single reference low-resolution image and multiple auxiliary low-quality frames within a diffusion model framework. By leveraging a multi-feature aggregation mechanism, FASR++ achieves identity-preserving face super-resolution without relying on explicit soft attribute annotations or gradient-based guidance. The method effectively exploits redundant information across video sequences and establishes state-of-the-art performance on standard benchmarks, simultaneously advancing both reconstruction quality (as measured by PSNR, SSIM, and LPIPS) and face recognition accuracy (in verification and identification tasks).

0 citationsRead paper

Rectifying Geometry-Induced Similarity Distortions for Real-World Aerial-Ground Person Re-Identification

Jan 29, 2026

This work addresses the degradation of cross-view similarity consistency caused by geometric distortions arising from extreme viewpoint and distance disparities between aerial and ground perspectives. To tackle this issue, the authors propose the Geometry-Induced Query-Key Transformation (GIQT) module, which explicitly corrects camera geometry–induced similarity distortions without altering feature representations or the underlying attention architecture. Leveraging a lightweight low-rank implementation and a geometry-conditioned prompting mechanism, GIQT incorporates a global view-adaptive prior, enabling the first explicit modeling and correction of geometric distortion within the attention mechanism. Evaluated on four aerial-ground person re-identification benchmarks, the method significantly enhances model robustness under extreme and unseen geometric conditions while incurring substantially lower computational overhead than existing approaches.

0 citationsRead paper

FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding

Jan 24, 2026

Current evaluation methods for video anomaly understanding (VAU) struggle to accurately assess models’ fine-grained descriptive capabilities regarding anomalous events, often diverging from human perception. This work reframes VAU as a tripartite parsing task—capturing the anomaly’s “What,” the involved entities “Who,” and the spatial context “Where”—and introduces FineW3, a new benchmark dataset, along with FVScore, a human-aligned evaluation metric. FVScore enables the first interpretable, fine-grained assessment of large vision-language models (LVLMs) based on their coverage of critical visual elements. Structured automatic annotations augment manual labeling, and scoring is grounded in key visual components. Human evaluations demonstrate that FVScore significantly outperforms existing metrics. Experiments further reveal that LVLMs underperform on tasks requiring spatiotemporal fine-grained reasoning but excel in scenarios with static or strong visual cues.

0 citationsRead paper
Recent publications

Latest Papers

CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device

Aug 10, 2026

This study addresses the inefficiency and low accuracy of manual length measurement for industrial cable reels by proposing an end-to-end vision-based method that estimates cable length from a single smartphone image. For the first time, camera calibration, instance segmentation, 6D pose estimation, and volume computation are integrated into a handheld device pipeline. The approach accommodates diverse reel types and cable diameters, enabling rapid length estimation from just one image. Trained on 1,000 annotated images, the instance segmentation model achieves a test-set mAP50 of 99.5%. Evaluated on 75 real-world cable reels, the system yields an average absolute percentage error of 4.90% with a per-image inference time of only 5.66 milliseconds, significantly enhancing both efficiency and precision in field measurements.

0 citationsRead paper

VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

Jul 07, 2026

This work addresses the lack of a unified evaluation protocol that hinders direct comparison among deepfake detection approaches across three dominant paradigms: commercial APIs, zero-shot vision-language models, and open-source detectors. To bridge this gap, the authors introduce VendorBench-100, a benchmark comprising a standardized corpus of 100 adversarial images, evaluated under consistent output formatting and a dual-metric strategy prioritizing the Matthews Correlation Coefficient (MCC) with ROC-AUC as a secondary measure. For the first time, 36 representative models are systematically assessed within a common framework. The results reveal that commercial APIs generally achieve the best performance, while certain open-source models rival top-tier vision-language models. Notably, high ROC-AUC scores do not necessarily correspond to high MCC values, underscoring the critical influence of metric selection on evaluation outcomes.

0 citationsRead paper

Robust Face Super-Resolution and Recognition Through Multi-Feature Aggregation in Diffusion Models

Jul 06, 2026

This work addresses the challenge of face super-resolution in surveillance scenarios, where low resolution, pose variations, uneven illumination, and occlusions severely hinder recognition performance. Conventional super-resolution methods often introduce identity distortions that degrade downstream tasks. To overcome this, the authors propose FASR++, a novel approach that, for the first time, integrates features from a single reference low-resolution image and multiple auxiliary low-quality frames within a diffusion model framework. By leveraging a multi-feature aggregation mechanism, FASR++ achieves identity-preserving face super-resolution without relying on explicit soft attribute annotations or gradient-based guidance. The method effectively exploits redundant information across video sequences and establishes state-of-the-art performance on standard benchmarks, simultaneously advancing both reconstruction quality (as measured by PSNR, SSIM, and LPIPS) and face recognition accuracy (in verification and identification tasks).

0 citationsRead paper

Rectifying Geometry-Induced Similarity Distortions for Real-World Aerial-Ground Person Re-Identification

Jan 29, 2026

This work addresses the degradation of cross-view similarity consistency caused by geometric distortions arising from extreme viewpoint and distance disparities between aerial and ground perspectives. To tackle this issue, the authors propose the Geometry-Induced Query-Key Transformation (GIQT) module, which explicitly corrects camera geometry–induced similarity distortions without altering feature representations or the underlying attention architecture. Leveraging a lightweight low-rank implementation and a geometry-conditioned prompting mechanism, GIQT incorporates a global view-adaptive prior, enabling the first explicit modeling and correction of geometric distortion within the attention mechanism. Evaluated on four aerial-ground person re-identification benchmarks, the method significantly enhances model robustness under extreme and unseen geometric conditions while incurring substantially lower computational overhead than existing approaches.

0 citationsRead paper

FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding

Jan 24, 2026

Current evaluation methods for video anomaly understanding (VAU) struggle to accurately assess models’ fine-grained descriptive capabilities regarding anomalous events, often diverging from human perception. This work reframes VAU as a tripartite parsing task—capturing the anomaly’s “What,” the involved entities “Who,” and the spatial context “Where”—and introduces FineW3, a new benchmark dataset, along with FVScore, a human-aligned evaluation metric. FVScore enables the first interpretable, fine-grained assessment of large vision-language models (LVLMs) based on their coverage of critical visual elements. Structured automatic annotations augment manual labeling, and scoring is grounded in key visual components. Human evaluations demonstrate that FVScore significantly outperforms existing metrics. Experiments further reveal that LVLMs underperform on tasks requiring spatiotemporal fine-grained reasoning but excel in scenarios with static or strong visual cues.

0 citationsRead paper