IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
研究通过IBBench-Light测试外部指令的任务条件响应,使用配对评估方法和PECA指标,检验了模型在执行和处理指令上的表现。
研究通过IBBench-Light测试外部指令的任务条件响应,使用配对评估方法和PECA指标,检验了模型在执行和处理指令上的表现。
This study addresses quality, efficiency, and cost bottlenecks in AI-assisted Linux and IoT malware analysis for 2024–2025. Methodologically, it introduces an “AI–expert collaborative” reverse-engineering paradigm: first systematically validating the practical boundaries of Claude 3.5/3.7 Sonnet on real-world malware; then integrating R2AI (an AI extension for Radare2), static/dynamic analysis, human-in-the-loop prompt engineering, and a feedback-closed loop—where domain experts provide real-time guidance to mitigate hallucination and cyclic reasoning. Results demonstrate: (1) analytical accuracy matching or exceeding manual analysis; (2) significantly accelerated end-to-end analysis—including error correction; and (3) per-sample cost substantially lower than a senior analyst’s daily rate. The core contribution is a deployable, empirically validated AI–human collaborative framework, rigorously demonstrated for state-of-the-art LLMs applied to complex binary reverse engineering—proving both technical efficacy and operational cost-efficiency.
研究通过IBBench-Light测试外部指令的任务条件响应,使用配对评估方法和PECA指标,检验了模型在执行和处理指令上的表现。
This study addresses quality, efficiency, and cost bottlenecks in AI-assisted Linux and IoT malware analysis for 2024–2025. Methodologically, it introduces an “AI–expert collaborative” reverse-engineering paradigm: first systematically validating the practical boundaries of Claude 3.5/3.7 Sonnet on real-world malware; then integrating R2AI (an AI extension for Radare2), static/dynamic analysis, human-in-the-loop prompt engineering, and a feedback-closed loop—where domain experts provide real-time guidance to mitigate hallucination and cyclic reasoning. Results demonstrate: (1) analytical accuracy matching or exceeding manual analysis; (2) significantly accelerated end-to-end analysis—including error correction; and (3) per-sample cost substantially lower than a senior analyst’s daily rate. The core contribution is a deployable, empirically validated AI–human collaborative framework, rigorously demonstrated for state-of-the-art LLMs applied to complex binary reverse engineering—proving both technical efficacy and operational cost-efficiency.