JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of structured data and evaluation benchmarks supporting the full paleographic exegesis pipeline, as existing methods primarily focus on ancient script recognition and retrieval. To bridge this gap, the authors propose the Ancient Chinese Character Exegesis (ACCE) visual question answering task and introduce JieZi, a dataset comprising over 500,000 expert-verified question-answer pairs. They further design a four-tier evaluation framework encompassing glyph recognition, structural analysis, semantic interpretation, and diachronic evolution. Data authenticity and assessment reliability are ensured through expert-guided templates, source-text-constrained generation, and multi-stage human verification. Experimental results demonstrate that while current multimodal large language models perform adequately on basic recognition, they exhibit significant deficiencies in deep semantic reasoning and diachronic understanding; fine-tuning on JieZi substantially improves performance across all evaluation tiers.
📝 Abstract
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis. To address this limitation, we introduce Ancient Chinese Character Exegesis (ACCE), a vision-language question answering (VQA) task that models the scholarly exegesis process. ACCE is organized into four progressive levels: basic character identification, glyph-form analysis, meaning exegesis, and diachronic evolution analysis. To support this task, we construct two complementary resources. JieZi-Dataset is the first large-scale, expert-audited VQA training dataset for ACCE, comprising over 500K QA pairs. It is constructed via a pipeline that reduces factual errors by constraining generation with expert-designed templates and source-text references. Human verification is further applied at each key stage to ensure scholarly accuracy. JieZi-Bench is an evaluation benchmark aligned with the exegesis process, constructed and verified by human experts to ensure evaluation reliability. It consists of four levels with reference answers curated from authoritative lexicographic works held separate from the training data. Experiments on multimodal large language models show that current models perform well on basic identification but struggle with glyph analysis, semantic reasoning, and diachronic understanding. Fine-tuning on JieZi-Dataset substantially improves performance across all four levels. Code and dataset are available at https://github.com/Ran00w/JieZi.
Problem

Research questions and friction points this paper is trying to address.

Ancient Chinese Character Exegesis
vision-language question answering
structured dataset
evaluation benchmark
scholarly analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ancient Chinese Character Exegesis
Vision-Language QA
Expert-Audited Dataset
Diachronic Evolution Analysis
Multimodal Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.