Survey on Publicly Available Sinhala Natural Language Processing Tools and Research
The Sinhala NLP community suffers from fragmented resources, absence of systematic surveys, and lack of collaborative benchmarks. Method: This paper introduces the first dynamically updated, open-source panoramic survey of Sinhala NLP. Leveraging bibliometric analysis, automated crawling and classification of open-source tools, multidimensional metadata annotation, and continuous arXiv tracking, it systematically catalogs dozens of global Sinhala NLP projects and tools. Contribution/Results: It identifies critical technical gaps and reuse pathways, and innovatively establishes a sustainably maintained knowledge graph and collaborative benchmark suite—filling a key void in unified surveys for low-resource language NLP. The survey significantly enhances community visibility, reproducibility, and interoperability, and has become the central reference and coordination hub for Sinhala NLP researchers in Sri Lanka and worldwide.