The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant challenge of efficiently identifying suicidal ideation and anti-suicidal求助 signals within the vast, noisy textual data of Russian-language social media—a task critical for timely intervention yet hindered by linguistic and contextual complexity. To this end, the authors propose a systematic methodology encompassing fine-grained category definitions, detailed annotation guidelines, and a multi-stage human annotation and validation pipeline. As a key contribution, they construct and publicly release the first large-scale Russian dataset of pre-suicidal and anti-suicidal signals, comprising over 50,000 annotated posts. The project also provides full code and documentation, and demonstrates the dataset’s utility through baseline classification models evaluated at multiple levels of granularity, thereby establishing a reproducible benchmark for future research on automated detection of suicide-related content in Russian.
📝 Abstract
The suicide is a terrifying act of a person who is misled by his own mental state. This problem arises across many countries. Sadly, Russia also has quite high number of persons who committed suicide. Luckily, a subset of these people writes their struggles in social media, allowing a way to find them and help. However, these valuable texts disappearing in many irrelevant texts which is considerably slowing down the decision process about person's suicidal risk. To tackle this problem, in this work we have presented a detailed methodology of building the dataset for detecting texts that describe presuicidal and anti-suicidal signals. This methodology describes the process of instruction and class table creation, the process of annotation, verification and post-annotation correction. Guiding by this methodology, we collect and annotate a large-scale Russian dataset with more than 50 thousand texts from social media. We provide a count statistic of the dataset as well as common problems in annotation. We also conduct basic experiments of building the classification models to show the on go performance on different levels of annotation. Furthermore, we make the dataset, code and all materials publicly available.
Problem

Research questions and friction points this paper is trying to address.

suicide detection
social media
presuicidal signals
anti-suicidal signals
Russian text
Innovation

Methods, ideas, or system contributions that make the work stand out.

suicide detection
large-scale dataset
social media text
annotation methodology
Russian NLP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
I
Igor Buyanov
Federal Research Center "Computer Science and Control" of the Russian Academy of Sciences, 44, build. 2, Vavilova St, Moscow, 119333, Russia.
D
Darya Yaskova
MTS AI, 23, build 5, Podsosenskiy lane, Moscow, 105062, Russia.
D
Danil Serenko
Federal Research Center "Computer Science and Control" of the Russian Academy of Sciences, 44, build. 2, Vavilova St, Moscow, 119333, Russia.
D
Danil Shkereda
Federal Research Center "Computer Science and Control" of the Russian Academy of Sciences, 44, build. 2, Vavilova St, Moscow, 119333, Russia.
A
Andrey Yaskov
Yandex, 16, Lev Tolstoy St, Moscow, 119021, Russia.
I
Ilya Sochenkov
Federal Research Center "Computer Science and Control" of the Russian Academy of Sciences, 44, build. 2, Vavilova St, Moscow, 119333, Russia.; Kharkevich Institute for Information Transmission Problems of the Russian Academy of Sciences, 19, build 1, Bolshoy Karetny Lane, Moscow, 127051, Russia.; Ivannikov Institute for System Programming of the Russian Academy of Sciences, 25, Alexander Solzhenitsyn St, Moscow, 109004, Russia.