Generating Attacks for LLMs with GFlowNets
Current red-teaming approaches for large language models (LLMs) rely heavily on manual efforts or static datasets, resulting in low efficiency and limited capacity to uncover deep-seated security vulnerabilities. This work proposes the first automatic and adaptive red-teaming framework based on Generative Flow Networks (GFlowNets), which leverages an attacker LLM to dynamically generate highly creative adversarial inputs. The framework autonomously identifies vulnerabilities in target models and quantifies their robustness without human intervention. By introducing GFlowNets into LLM red-teaming for the first time, the method outperforms existing benchmarks in English attack generation and pioneers support for automatic adversarial input generation in low-resource languages such as Turkish, substantially enhancing test coverage and evaluation efficiency.