A Word-Level Digital Reader of the Prasthanatrayi with Sankara's Bhasya: Corpus, Method, and an Open, Offline Reading Aid for the Advaita Vedanta Canon

📅 2026-07-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses word-level readability barriers in the *Triple Canon* and Śaṅkara’s commentary—arising from sandhi, compound formation, and dense scholarly prose—by presenting the first offline, open-source, word-level interactive reading system covering the complete text. The system integrates a rule-based sandhi splitter, an inflectional lexicon, corpus-based lookup tables, and a large language model, enhanced by an adversarial two-pass validation protocol and a human-in-the-loop correction mechanism. It encompasses 13 commentary units, 36,881 root-text tokens, and 95,587 surface forms from the commentary, achieving over 99% agreement with authoritative dictionaries at high-confidence analysis levels. This substantially enhances the readability and searchability of Sanskrit philosophical texts.
📝 Abstract
The Prasthanatrayi -- the ten principal Upanisads, the Brahmasutra, and the Bhagavadgita, with Sankara's commentaries (bhasya) -- is the foundational corpus of Advaita Vedanta. Continuous euphonic combination (sandhi), long compounds (samasa), and dense scholastic prose make it hard to read at the word level: where one word ends, and what each word means grammatically, are both obscured. We present an open, fully offline, word-level digital reader of the entire Prasthanatrayi with Sankara's bhasya. Every word -- of both the root text (mula) and the commentary -- is clickable and resolves to a pop-up giving its split (padaccheda), morphological analysis, and gloss. Because every word carries a lemma, the reader also acts as a concordance: a search on a dictionary headword retrieves all of that word's inflected and sandhi-hidden occurrences, and its occurrences inside compounds, across both layers. The resource covers thirteen commentarial units (2,971 verses, sutras, and prose sections; 36,881 analysed word-occurrences of root text) and a global dictionary of 95,587 distinct commentarial surface forms. We describe the corpus, the hybrid pipeline -- a rule-based sandhi splitter over an inflected-form lexicon and attested-corpus look-ups, with LLM-assisted analysis under an adversarial two-pass verification protocol -- and a durable human-review loop whose corrections survive every regeneration. An intrinsic evaluation against independent Sanskrit resources finds high-confidence analyses agree with an authoritative inflectional lexicon on over 99% of attested forms, and a band-blind adjudication confirms that quality degrades predictably across confidence bands, with errors concentrated in the low-confidence tier the review loop targets. The reader is a single self-contained HTML file needing no server or network, offered as a freely redistributable teaching and reading aid.
Problem

Research questions and friction points this paper is trying to address.

sandhi
samasa
word-level readability
Advaita Vedanta
morphological analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

word-level analysis
sandhi splitting
morphological parsing
offline digital reader
LLM-assisted verification
🔎 Similar Papers
No similar papers found.
T
Tamal Maharaj
Department of Computer Science, Ramakrishna Mission Vivekananda Educational and Research Institute, Belur Math, Howrah, West Bengal, India