Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems
This work addresses the challenges of language diversity, document heterogeneity, and deployment efficiency in OCR systems for multilingual and multimodal Indian government documents. The authors propose the Chitrapathak family of methods, which combines end-to-end training of a general-purpose vision-language model with efficient fine-tuning of pretrained OCR models for non-target languages. Additionally, they introduce Parichay, the first dedicated structured information extraction model tailored to nine categories of Indian governmental documents. Experimental results demonstrate that Chitrapathak-2 achieves state-of-the-art performance on Telugu (6.69 character-level ANLS) and ranks second on other languages, while offering 3–6× faster inference. Parichay attains an Exact Match score of 89.8% on key field extraction from government documents, demonstrating both high accuracy and computational efficiency.