Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建古典泰米尔语诗-注释对语料库,使用多种深度学习模型探索信息表示学习,以解决文本内容和结构的自动理解与生成问题。
📝 Abstract
We construct a corpus of 1,262 verse-commentary (urai) pairs from five Classical Tamil source sections, ranging from technical grammatical prose to modern paraphrase, and ask what information representation learning can recover. We train recurrent and Transformer encoders, a Siamese-style pair-matching network, an mBART-style encoder-decoder, and a decoder-only language model. Each analysis is interpreted against an appropriate control on the same data. TF-IDF provides a strong no-training lexical retrieval baseline, alongside representation analyses and generation controls for the learned models. A fixed string containing the 25 most frequent commentary words scores higher on generation overlap than the decoder-only model. Canonical correlation reaches 1.000 on Gaussian noise at these sample sizes, token-F1 spans only about 0.02-0.20 on this corpus, and the encoder-decoder continues to lower training loss for sixteen epochs after validation loss has begun to rise. One narrow result remains: the decoder-only model prefers authentic word order in 107 of 112 minimal-pair comparisons (95.5%), but does not reproduce held-out commentary content. We release the extraction and evaluation protocol; redistribution of the source commentaries remains subject to permission.
Problem

Research questions and friction points this paper is trying to address.

representation learning
Classical Tamil
verse-commentary pairs
Innovation

Methods, ideas, or system contributions that make the work stand out.

representation learning
classical tamil
verse-commentary pairs
decoder-only model
tf-idf
🔎 Similar Papers
No similar papers found.
A
Amrit Gopinath
Sri Sivasubramaniya Nadar College of Engineering, Chennai, India
S
Sangeetha Sivanesan
National Institute of Technology Tiruchirappalli, Tamil Nadu, India