RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
RiskChainBench通过合成数据和模型恢复被混淆的平台消息,并结合证据驱动的网页调查,解决网络滥用活动中的重定向和风险评估问题。
📝 Abstract
Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages separately, obscuring how target recovery affects downstream evidence acquisition. We introduce RiskChainBench, pairing 3,600 synthetic token-text restoration inputs from 600 source sessions with 600 corresponding human-labeled local web environments. A model first restores the message, operational intent, and destination; the same underlying model then acts as a VLM-driven web agent that investigates the correctly associated website and produces a frozen, evidence-cited risk report without message-side semantics or domain-reputation cues. We score restoration and correct-routing web investigation separately and compose them offline by applying the frozen primary-entry prediction as a gate to the same Task 2 result. Human labels determine task correctness, while a fixed multimodal evidence judge assesses faithfulness, sufficiency, completeness, and consistency. Across ten models, Entry Top-1 ranges from 35.2% to 95.2% and web decision accuracy from 26.3% to 62.8%; the leading systems differ across entry recovery, full reconstruction, website decisions, and fine-grained typing. Execution failures account for 31.9% of web runs, whereas post-decision type errors account for only 0.9%, identifying stable exploration and risk judgment as the principal bottlenecks. We release the benchmark, protocol, and resettable local sandbox.
Problem

Research questions and friction points this paper is trying to address.

Platform Abuse
Obfuscated Text
Risky Webpages
Evidence Acquisition
Redirection Instructions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Obfuscated Text Restoration
Web Investigation
Evidence-Grounded Report
Single Model Approach
Local Sandbox
Z
ZhuoXin Liu
Baidu; SmartFlowAI
Z
Zhiming Ma
SmartFlowAI; Northeastern University
Y
Ying Zhang
Baidu
M
Mengzheng Yang
People’s Public Security University of China
Y
Yifan Wang
People’s Public Security University of China
Z
Zhengqi Huang
Northeastern University
Y
Yanhan Zhou
Tsinghua University
Z
Zekun Lin
Baidu
J
Jun Zhang
Baidu
S
Shun Zhang
People’s Public Security University of China
Y
Yue Chen
JD Technology
Q
Qiao Zhao
Baidu; Tsinghua University
Peng Chen
Peng Chen
Ph.D. student, East China Normal University
Time Series Forecasting,LLM, Foundation Models