Institution profile

Toloka AI

Industry researcheurope · ru
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning

Apr 23, 2026

This study addresses the challenge of modeling similarity between narrative stories and learning effective narrative representations. The authors propose a triplet-based binary classification task grounded in narrative theory and human intuition: given an anchor story, the model determines which of two candidate stories is more similar to it. To support this task, they construct a high-quality dataset of human-annotated narrative triplets. Their approach integrates large language model (LLM) ensembles, fine-tuned pretrained embeddings, and pre- and post-processing strategies. In an evaluation involving 46 teams and 71 submissions, the LLM ensemble achieved the best performance on the classification task, while fine-tuning and preprocessing yielded comparable results for the embedding task. These findings suggest that automated narrative understanding still has considerable room for improvement.

0 citationsRead paper

Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies

Dec 16, 2025

This work addresses the dual challenges of fairness and practicality in NLP for multilingual and low-resource languages—particularly underrepresented ones—amid data scarcity and cultural heterogeneity. To this end, we propose a culturally adaptive, end-to-end NLP development paradigm. Methodologically, it integrates community-engaged data collection, self-supervised parallel sentence mining, few-shot machine translation fine-tuning, zero-shot text classification, and multimodal reasoning interfaces, all within a lightweight modeling framework. We present the first systematic consolidation of over ten cross-linguistic, multi-regional language case studies—including severely under-resourced varieties—packaged as an open-source toolkit and pedagogical resource. Our approach substantially lowers the technical barrier for low-resource language NLP development, enabling reproducible and scalable NLP applications across more than a dozen languages. The work advances equitable, sustainable, and community-driven language technology.

0 citationsRead paper

Surveying Professional Writers on AI: Limitations, Expectations, and Fears

Apr 07, 2025

This study investigates adoption barriers and impacts of AI writing tools—particularly large language models (LLMs)—in multilingual professional writing. Addressing a critical gap, it examines challenges faced by non-English writers through a mixed-methods design: a structured survey (N=301) with global professional writers across 25+ languages, in-depth interactive writing tasks (N=36), and cross-lingual textual analysis with thematic modeling. The study provides the first empirical evidence identifying key adoption impediments—including uneven linguistic support, weak domain adaptation, and stylistic homogenization risks—and proposes two novel evaluation dimensions: *stylistic adaptability* and *misinformation controllability*. It further identifies high-priority functional requirements: factual verification, register control, and iterative human-AI collaboration. Collectively, these findings offer empirically grounded guidance for developing ethically responsible, linguistically inclusive, and creator-centered LLM writing tools.

0 citationsRead paper

JEEM: Vision-Language Understanding in Four Arabic Dialects

Mar 27, 2025

Existing vision-language models (VLMs) exhibit poor cross-dialect generalization and weak cultural element comprehension when applied to Arabic visual understanding across major dialects (Jordanian, Emirati, Egyptian, Moroccan). Method: We introduce JEEM, the first multi-dialect Arabic vision-language evaluation benchmark, covering image captioning and visual question answering tasks with emphasis on cultural diversity and regional adaptation. JEEM is built upon manually annotated, culturally rich, multi-regional image data and employs a standardized protocol to evaluate five open-source Arabic VLMs alongside GPT-4V. Contribution/Results: Experiments reveal that all open-source models substantially underperform GPT-4V; though GPT-4V shows uneven dialect proficiency and limitations in visual reasoning, it remains the state-of-the-art. This work provides the first systematic analysis of cultural perception bottlenecks in multi-dialect Arabic VLMs, establishing a critical benchmark and empirical foundation for developing culturally aware Arabic VLMs.

0 citationsRead paper

REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities

Mar 17, 2025

This work investigates the effectiveness of large language models (LLMs) as evaluators for Russian-language generation, revealing a substantial performance gap compared to their English-language counterparts. To address this gap, we introduce REPA—the first fine-grained Russian error annotation dataset—comprising 1k queries and 2k responses annotated across 10 linguistic error types, alongside the first human-curated multidimensional preference annotation framework. We systematically evaluate six generative models and eight LLM-as-judge evaluators via human ranking, zero- and few-shot LLM-based evaluation, and bias analysis. Results indicate that LLM judges exhibit limited fine-grained discrimination capability in Russian, with low human–LLM preference alignment; positional and length biases significantly impair judgment consistency. This study fills a critical gap in Russian LLM evaluation and establishes a foundational benchmark dataset and methodology for non-English LLM assessment.

0 citationsRead paper
Recent publications

Latest Papers

SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning

Apr 23, 2026

This study addresses the challenge of modeling similarity between narrative stories and learning effective narrative representations. The authors propose a triplet-based binary classification task grounded in narrative theory and human intuition: given an anchor story, the model determines which of two candidate stories is more similar to it. To support this task, they construct a high-quality dataset of human-annotated narrative triplets. Their approach integrates large language model (LLM) ensembles, fine-tuned pretrained embeddings, and pre- and post-processing strategies. In an evaluation involving 46 teams and 71 submissions, the LLM ensemble achieved the best performance on the classification task, while fine-tuning and preprocessing yielded comparable results for the embedding task. These findings suggest that automated narrative understanding still has considerable room for improvement.

0 citationsRead paper

Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies

Dec 16, 2025

This work addresses the dual challenges of fairness and practicality in NLP for multilingual and low-resource languages—particularly underrepresented ones—amid data scarcity and cultural heterogeneity. To this end, we propose a culturally adaptive, end-to-end NLP development paradigm. Methodologically, it integrates community-engaged data collection, self-supervised parallel sentence mining, few-shot machine translation fine-tuning, zero-shot text classification, and multimodal reasoning interfaces, all within a lightweight modeling framework. We present the first systematic consolidation of over ten cross-linguistic, multi-regional language case studies—including severely under-resourced varieties—packaged as an open-source toolkit and pedagogical resource. Our approach substantially lowers the technical barrier for low-resource language NLP development, enabling reproducible and scalable NLP applications across more than a dozen languages. The work advances equitable, sustainable, and community-driven language technology.

0 citationsRead paper

Surveying Professional Writers on AI: Limitations, Expectations, and Fears

Apr 07, 2025

This study investigates adoption barriers and impacts of AI writing tools—particularly large language models (LLMs)—in multilingual professional writing. Addressing a critical gap, it examines challenges faced by non-English writers through a mixed-methods design: a structured survey (N=301) with global professional writers across 25+ languages, in-depth interactive writing tasks (N=36), and cross-lingual textual analysis with thematic modeling. The study provides the first empirical evidence identifying key adoption impediments—including uneven linguistic support, weak domain adaptation, and stylistic homogenization risks—and proposes two novel evaluation dimensions: *stylistic adaptability* and *misinformation controllability*. It further identifies high-priority functional requirements: factual verification, register control, and iterative human-AI collaboration. Collectively, these findings offer empirically grounded guidance for developing ethically responsible, linguistically inclusive, and creator-centered LLM writing tools.

0 citationsRead paper

JEEM: Vision-Language Understanding in Four Arabic Dialects

Mar 27, 2025

Existing vision-language models (VLMs) exhibit poor cross-dialect generalization and weak cultural element comprehension when applied to Arabic visual understanding across major dialects (Jordanian, Emirati, Egyptian, Moroccan). Method: We introduce JEEM, the first multi-dialect Arabic vision-language evaluation benchmark, covering image captioning and visual question answering tasks with emphasis on cultural diversity and regional adaptation. JEEM is built upon manually annotated, culturally rich, multi-regional image data and employs a standardized protocol to evaluate five open-source Arabic VLMs alongside GPT-4V. Contribution/Results: Experiments reveal that all open-source models substantially underperform GPT-4V; though GPT-4V shows uneven dialect proficiency and limitations in visual reasoning, it remains the state-of-the-art. This work provides the first systematic analysis of cultural perception bottlenecks in multi-dialect Arabic VLMs, establishing a critical benchmark and empirical foundation for developing culturally aware Arabic VLMs.

0 citationsRead paper

REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities

Mar 17, 2025

This work investigates the effectiveness of large language models (LLMs) as evaluators for Russian-language generation, revealing a substantial performance gap compared to their English-language counterparts. To address this gap, we introduce REPA—the first fine-grained Russian error annotation dataset—comprising 1k queries and 2k responses annotated across 10 linguistic error types, alongside the first human-curated multidimensional preference annotation framework. We systematically evaluate six generative models and eight LLM-as-judge evaluators via human ranking, zero- and few-shot LLM-based evaluation, and bias analysis. Results indicate that LLM judges exhibit limited fine-grained discrimination capability in Russian, with low human–LLM preference alignment; positional and length biases significantly impair judgment consistency. This study fills a critical gap in Russian LLM evaluation and establishes a foundational benchmark dataset and methodology for non-English LLM assessment.

0 citationsRead paper