Institution profile

Hacettepe University

Academic institutioneurope · tr
Official website
Research library48linked papers
Opportunities0open roles
Selected work

Representative Papers

Distributed Intrusion Detection in Dynamic Networks of UAVs using Few-Shot Federated Learning

Jan 22, 2025

Addressing the challenges of distributed intrusion detection in high-speed, dynamic Flying Ad-hoc Networks (FANETs)—characterized by constrained communication bandwidth, privacy sensitivity, stringent energy limitations, and frequent link disruptions—this paper proposes a novel framework integrating Few-Shot Learning (FSL) with Federated Learning (FL). Uniquely embedding FSL into the FL training pipeline, our approach drastically reduces dependency on labeled data at both local and global levels, cutting required training samples for routing attack detection by over 60%. It simultaneously ensures end-device privacy preservation, ultra-low-power operation, and robustness against packet loss. Experimental results demonstrate significant reductions in communication overhead and computational latency, alongside extended UAV battery lifetime. The framework establishes an efficient, secure, and sustainable paradigm for distributed intrusion detection in resource-constrained, highly dynamic edge networks.

1 citationsRead paper

Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

Aug 11, 2026

This work addresses key challenges in resource-constrained sign language research—namely, the absence of dense lexical annotations, loose alignment between spoken and signed content in broadcast news, and fragmented pseudo-labels caused by the morphological complexity of languages like Turkish. To overcome these issues, the authors propose a weakly supervised pretraining approach that requires no manual gloss annotations. By integrating rule-based lemmatization with large language model (LLM)-constrained lexical normalization, they generate high-quality pseudo-gloss labels and pretrain a transferable sign language encoder on a newly curated TSL-News corpus. This study presents the first application of transcription-based, gloss-free pretraining to cross-dataset sign language localization, achieving substantial performance gains: the top-5 mean IoU improves from 0.235 to 0.465 (with 56.2% of samples attaining IoU ≥ 0.5), while downstream translation scores reach 11.04 BLEU-4 and 27.43 ROUGE.

0 citationsRead paper

ODE-Based Transformer Decoders for Iterative Sign Language Translation

Aug 11, 2026

This work addresses the high computational cost and low parameter efficiency of current sign language translation models, which often rely on scaling up model size for performance gains. The authors propose a novel reconstruction of the Transformer decoder from the perspective of ordinary differential equations (ODEs), introducing higher-order numerical integration schemes—such as Runge-Kutta methods RK-2 and RK-4—into sign language translation for the first time. This approach replaces conventional residual connections with more accurate and stable iterative optimization, enhancing model expressiveness without increasing parameter count. Experimental results demonstrate that the proposed method achieves BLEU-4 scores of 22.96 and 19.34 on the PHOENIX-2014-T and CSL-Daily datasets, respectively, outperforming baseline models while using fewer decoder layers and iteration steps.

0 citationsRead paper

EeveeDark: A Binary Neural Framework for Low-Light Video Enhancement via Event-Guided Sensor-Level Fusion

Jul 07, 2026

This work addresses the challenge of achieving high-quality video enhancement under extreme low-light conditions while maintaining computational efficiency in resource-constrained settings. To this end, it introduces binary neural networks (BNNs) into RAW-event multimodal fusion for the first time, proposing a modality-specific binary encoder, a lightweight cross-modal fusion module, and an event-guided skip gating mechanism to enable dynamic spatiotemporal optimization. Evaluated on both synthetic and real-world low-light datasets, the proposed method significantly outperforms existing BNN-based approaches, delivering superior enhancement quality while substantially reducing computational overhead. This approach effectively strikes a balance between performance and efficiency, making it particularly suitable for practical deployment in low-power or embedded vision systems.

0 citationsRead paper
Recent publications

Latest Papers

Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

Aug 11, 2026

This work addresses key challenges in resource-constrained sign language research—namely, the absence of dense lexical annotations, loose alignment between spoken and signed content in broadcast news, and fragmented pseudo-labels caused by the morphological complexity of languages like Turkish. To overcome these issues, the authors propose a weakly supervised pretraining approach that requires no manual gloss annotations. By integrating rule-based lemmatization with large language model (LLM)-constrained lexical normalization, they generate high-quality pseudo-gloss labels and pretrain a transferable sign language encoder on a newly curated TSL-News corpus. This study presents the first application of transcription-based, gloss-free pretraining to cross-dataset sign language localization, achieving substantial performance gains: the top-5 mean IoU improves from 0.235 to 0.465 (with 56.2% of samples attaining IoU ≥ 0.5), while downstream translation scores reach 11.04 BLEU-4 and 27.43 ROUGE.

0 citationsRead paper

ODE-Based Transformer Decoders for Iterative Sign Language Translation

Aug 11, 2026

This work addresses the high computational cost and low parameter efficiency of current sign language translation models, which often rely on scaling up model size for performance gains. The authors propose a novel reconstruction of the Transformer decoder from the perspective of ordinary differential equations (ODEs), introducing higher-order numerical integration schemes—such as Runge-Kutta methods RK-2 and RK-4—into sign language translation for the first time. This approach replaces conventional residual connections with more accurate and stable iterative optimization, enhancing model expressiveness without increasing parameter count. Experimental results demonstrate that the proposed method achieves BLEU-4 scores of 22.96 and 19.34 on the PHOENIX-2014-T and CSL-Daily datasets, respectively, outperforming baseline models while using fewer decoder layers and iteration steps.

0 citationsRead paper

EeveeDark: A Binary Neural Framework for Low-Light Video Enhancement via Event-Guided Sensor-Level Fusion

Jul 07, 2026

This work addresses the challenge of achieving high-quality video enhancement under extreme low-light conditions while maintaining computational efficiency in resource-constrained settings. To this end, it introduces binary neural networks (BNNs) into RAW-event multimodal fusion for the first time, proposing a modality-specific binary encoder, a lightweight cross-modal fusion module, and an event-guided skip gating mechanism to enable dynamic spatiotemporal optimization. Evaluated on both synthetic and real-world low-light datasets, the proposed method significantly outperforms existing BNN-based approaches, delivering superior enhancement quality while substantially reducing computational overhead. This approach effectively strikes a balance between performance and efficiency, making it particularly suitable for practical deployment in low-power or embedded vision systems.

0 citationsRead paper

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

Jun 01, 2026

Existing generative video models struggle to achieve human-centric camera control aligned with cinematic language, often producing random camera trajectories, spatial inconsistencies, and insufficient focus on the human subject. This work proposes a human-centric camera parameterization method that formalizes cinematic composition principles into computable, human-relative camera parameters for the first time. It introduces a domain-specific language (DSL) that coordinates with a multimodal large language model to map natural language instructions and human motion into cinematic keyframe shots, followed by deterministic interpolation to generate smooth, continuous camera trajectories. Evaluated on a newly curated dataset of 34K text–motion–camera aligned samples, the approach significantly outperforms existing methods on composition-oriented metrics, enabling controllable and aesthetically cinematic human-centric video generation.

0 citationsRead paper