Institution profile

Wikimedia Foundation

Academic institutionnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Fueling Volunteer Growth: the case of Wikipedia Administrators

Jan 27, 2026

This study addresses the persistent decline in administrator numbers across highly active Wikipedia language editions—a trend that threatens the platform’s long-term sustainability. Integrating large-scale log data from 284 language editions since 2018, over 3,000 survey responses, and 12 in-depth interviews, this work reveals for the first time that administrator attrition stems primarily from insufficient recruitment rather than abnormal departures. Key barriers include low awareness of administrative roles, ambiguous eligibility criteria, and overly stringent onboarding procedures. Although a majority of language editions exhibit net growth in administrators, two-thirds of highly active editions show a declining trend. Building on these findings, the study proposes actionable strategies to optimize recruitment processes, offering empirical evidence and practical pathways to strengthen Wikipedia’s governance mechanisms and ensure their sustainability.

0 citationsRead paper

Question-to-Question Retrieval for Hallucination-Free Knowledge Access: An Approach for Wikipedia and Wikidata Question Answering

Jan 20, 2025

Large-scale knowledge bases (e.g., Wikipedia/Wikidata) suffer from hallucination and inefficiency in question answering. Method: We propose a “question-question matching” retrieval paradigm: instruction-tuned LLMs (e.g., Llama-3) generate multi-perspective questions for each knowledge unit; these questions are embedded into a dense vector space using Sentence-BERT or ColBERT. User queries are matched directly against the precomputed question index—enabling zero-shot, generation-free, semantically aligned knowledge access. Crucially, this approach replaces document-level retrieval with question-level retrieval and integrates Wikidata’s RDF schema for structured fact mapping. Contributions/Results: Experiments on Wikipedia and Wikidata achieve >90% top-1 accuracy, sub-100ms latency, and support multimodal (text + multimedia) QA. The method significantly improves scalability, reliability, and retrieval precision while eliminating LLM hallucination.

0 citationsRead paper
Recent publications

Latest Papers

Fueling Volunteer Growth: the case of Wikipedia Administrators

Jan 27, 2026

This study addresses the persistent decline in administrator numbers across highly active Wikipedia language editions—a trend that threatens the platform’s long-term sustainability. Integrating large-scale log data from 284 language editions since 2018, over 3,000 survey responses, and 12 in-depth interviews, this work reveals for the first time that administrator attrition stems primarily from insufficient recruitment rather than abnormal departures. Key barriers include low awareness of administrative roles, ambiguous eligibility criteria, and overly stringent onboarding procedures. Although a majority of language editions exhibit net growth in administrators, two-thirds of highly active editions show a declining trend. Building on these findings, the study proposes actionable strategies to optimize recruitment processes, offering empirical evidence and practical pathways to strengthen Wikipedia’s governance mechanisms and ensure their sustainability.

0 citationsRead paper

Question-to-Question Retrieval for Hallucination-Free Knowledge Access: An Approach for Wikipedia and Wikidata Question Answering

Jan 20, 2025

Large-scale knowledge bases (e.g., Wikipedia/Wikidata) suffer from hallucination and inefficiency in question answering. Method: We propose a “question-question matching” retrieval paradigm: instruction-tuned LLMs (e.g., Llama-3) generate multi-perspective questions for each knowledge unit; these questions are embedded into a dense vector space using Sentence-BERT or ColBERT. User queries are matched directly against the precomputed question index—enabling zero-shot, generation-free, semantically aligned knowledge access. Crucially, this approach replaces document-level retrieval with question-level retrieval and integrates Wikidata’s RDF schema for structured fact mapping. Contributions/Results: Experiments on Wikipedia and Wikidata achieve >90% top-1 accuracy, sub-100ms latency, and support multimodal (text + multimedia) QA. The method significantly improves scalability, reliability, and retrieval precision while eliminating LLM hallucination.

0 citationsRead paper