The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian
This study addresses the significant challenge of efficiently identifying suicidal ideation and anti-suicidal求助 signals within the vast, noisy textual data of Russian-language social media—a task critical for timely intervention yet hindered by linguistic and contextual complexity. To this end, the authors propose a systematic methodology encompassing fine-grained category definitions, detailed annotation guidelines, and a multi-stage human annotation and validation pipeline. As a key contribution, they construct and publicly release the first large-scale Russian dataset of pre-suicidal and anti-suicidal signals, comprising over 50,000 annotated posts. The project also provides full code and documentation, and demonstrates the dataset’s utility through baseline classification models evaluated at multiple levels of granularity, thereby establishing a reproducible benchmark for future research on automated detection of suicide-related content in Russian.