PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models
Static benchmarking inadequately captures LLM security vulnerabilities exposed in real-world online community practices. To address this, we propose a community-driven dynamic monitoring paradigm, focusing on psychological attacks as the primary threat vector, revealing that capability advancement and safety improvement are misaligned. Methodologically, we design a cross-platform data collection system, a multidimensional scoring framework, and a scalable monitoring architecture—integrating horizontal comparative analysis with fine-grained vulnerability classification. Over five months, we conduct an empirical study across nine commercial LLMs. Our approach identifies 198 novel vulnerabilities with 78% classification accuracy; psychological attacks exhibit significantly higher detection rates than traditional exploit-based techniques yet demonstrate low cross-model transferability—highlighting their stealthiness and model specificity. This work pioneers systematic, continuous discovery and quantitative evaluation of community-emergent LLM vulnerabilities.