Institution profile

Hishab Singapore Pte. Ltd

Industry researchasia · sg
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration

Nov 27, 2025

Existing multilingual models exhibit limited robustness for non-standard Romanized Hindi and Bengali text prevalent in South Asian social media, particularly with respect to phonetic variation, orthographic diversity, code-mixing, and low-resource adaptation. To address this, we introduce the first large-scale, high-diversity parallel dataset for Romanized-to-native script transliteration—comprising 1.8 million Hindi and 1.0 million Bengali sentence pairs—carefully curated to cover extensive phonological and orthographic variation. Leveraging the MarianMT framework, we train a multilingual sequence-to-sequence model specifically optimized for this task. Our approach significantly improves transliteration robustness in low-resource and code-mixed settings, outperforming state-of-the-art multilingual models across both BLEU and Character Error Rate (CER) metrics. This work bridges a critical gap in both high-quality transliteration resources and modeling capabilities for Romanized South Asian languages.

0 citationsRead paper
Recent publications

Latest Papers

Modeling Romanized Hindi and Bengali: Dataset Creation and Multilingual LLM Integration

Nov 27, 2025

Existing multilingual models exhibit limited robustness for non-standard Romanized Hindi and Bengali text prevalent in South Asian social media, particularly with respect to phonetic variation, orthographic diversity, code-mixing, and low-resource adaptation. To address this, we introduce the first large-scale, high-diversity parallel dataset for Romanized-to-native script transliteration—comprising 1.8 million Hindi and 1.0 million Bengali sentence pairs—carefully curated to cover extensive phonological and orthographic variation. Leveraging the MarianMT framework, we train a multilingual sequence-to-sequence model specifically optimized for this task. Our approach significantly improves transliteration robustness in low-resource and code-mixed settings, outperforming state-of-the-art multilingual models across both BLEU and Character Error Rate (CER) metrics. This work bridges a critical gap in both high-quality transliteration resources and modeling capabilities for Romanized South Asian languages.

0 citationsRead paper