PhishSSL: Self-Supervised Contrastive Learning for Phishing Website Detection

📅 2025-10-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Phishing website detection faces challenges of high annotation costs and poor adaptability to emerging attacks. To address these, this paper proposes an unsupervised self-supervised contrastive learning framework that effectively identifies malicious websites without human-labeled data. Methodologically, it innovatively integrates hybrid tabular data augmentation with an adaptive feature attention mechanism, generating semantically consistent augmented views while enhancing discriminative feature representation. Evaluated on three heterogeneous datasets, the framework consistently outperforms state-of-the-art unsupervised and self-supervised methods. Ablation studies validate the individual contributions of each component. The approach demonstrates strong generalization capability, cross-dataset transferability, and robust stability under varying conditions. Overall, it establishes a novel paradigm for phishing detection in low-resource settings, advancing practical deployment where labeled data is scarce or costly to obtain.

Technology Category

Application Category

📝 Abstract
Phishing websites remain a persistent cybersecurity threat by mimicking legitimate sites to steal sensitive user information. Existing machine learning-based detection methods often rely on supervised learning with labeled data, which not only incurs substantial annotation costs but also limits adaptability to novel attack patterns. To address these challenges, we propose PhishSSL, a self-supervised contrastive learning framework that eliminates the need for labeled phishing data during training. PhishSSL combines hybrid tabular augmentation with adaptive feature attention to produce semantically consistent views and emphasize discriminative attributes. We evaluate PhishSSL on three phishing datasets with distinct feature compositions. Across all datasets, PhishSSL consistently outperforms unsupervised and self-supervised baselines, while ablation studies confirm the contribution of each component. Moreover, PhishSSL maintains robust performance despite the diversity of feature sets, highlighting its strong generalization and transferability. These results demonstrate that PhishSSL offers a promising solution for phishing website detection, particularly effective against evolving threats in dynamic Web environments.
Problem

Research questions and friction points this paper is trying to address.

Detecting phishing websites without labeled training data
Addressing high annotation costs and limited attack adaptability
Improving generalization across diverse website feature sets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised contrastive learning eliminates labeled data need
Hybrid tabular augmentation creates semantically consistent views
Adaptive feature attention emphasizes discriminative phishing attributes
🔎 Similar Papers
No similar papers found.
W
Wenhao Li
Cybersecurity Research Centre, Universiti Sains Malaysia, Pulau Pinang, Malaysia
S
Selvakumar Manickam
Cybersecurity Research Centre, Universiti Sains Malaysia, Pulau Pinang, Malaysia
Y
Yung-Wey Chong
School of Computer Sciences, Universiti Sains Malaysia, Pulau Pinang, Malaysia
Shankar Karuppayah
Shankar Karuppayah
Deputy Director and Senior Lecturer at Universiti Sains Malaysia
BotnetsCyber SecurityPeer-to-peer networks
Priyadarsi Nanda
Priyadarsi Nanda
Faculty of Engineering and IT, University of Technology Sydney, Sydney, Australia
B
Binyong Li
Chengdu University of Information Technology, Chengdu, China