A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming

📅 2025-05-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Public WebShell malware datasets are scarce, and existing generation methods suffer from low payload diversity and high redundancy. Method: This paper proposes the first reward-driven, automated WebShell generation framework tailored for red-team operations. It introduces a novel LLM-based token standardization modeling system covering seven distinct obfuscation patterns and integrates supervised fine-tuning (SFT) with proximal policy optimization (PPO) for preference alignment—using malicious samples as “chosen” and benign code as “rejected.” Contribution/Results: The generated WebShells comprehensively cover all seven obfuscation patterns. Empirical evaluation shows a 32.7% average improvement in evasion success rate against mainstream detectors and a 61.4% reduction in redundancy, significantly outperforming baseline approaches.

Technology Category

Application Category

📝 Abstract
Frequent cyber-attacks have elevated WebShell exploitation and defense to a critical research focus within network security. However, there remains a significant shortage of publicly available, well-categorized malicious-code datasets organized by obfuscation method. Existing malicious-code generation methods, which primarily rely on prompt engineering, often suffer from limited diversity and high redundancy in the payloads they produce. To address these limitations, we propose extbf{RAWG}, a extbf{R}eward-driven extbf{A}utomated extbf{W}ebshell Malicious-code extbf{G}enerator designed for red-teaming applications. Our approach begins by categorizing webshell samples from common datasets into seven distinct types of obfuscation. We then employ a large language model (LLM) to extract and normalize key tokens from each sample, creating a standardized, high-quality corpus. Using this curated dataset, we perform supervised fine-tuning (SFT) on an open-source large model to enable the generation of diverse, highly obfuscated webshell malicious payloads. To further enhance generation quality, we apply Proximal Policy Optimization (PPO), treating malicious-code samples as"chosen"data and benign code as"rejected"data during reinforcement learning. Extensive experiments demonstrate that RAWG significantly outperforms current state-of-the-art methods in both payload diversity and escape effectiveness.
Problem

Research questions and friction points this paper is trying to address.

Lack of diverse, well-categorized webshell malicious-code datasets
Limited payload diversity in existing prompt-based generation methods
Need for automated generation of highly obfuscated webshell payloads
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward-driven automated webshell generator RAWG
LLM extracts tokens for supervised fine-tuning
PPO enhances generation with chosen-rejected data
🔎 Similar Papers
No similar papers found.