Unsupervised Domain Adaptation for Symbol Spotting in Historical Encrypted Manuscripts

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种无监督领域适应方法,通过三阶段流程解决历史加密手稿中符号识别问题,无需手动标注,并在实验中显著优于现有模型。
📝 Abstract
The decipherment of historical encrypted manuscripts poses a fundamental challenge in Digital Humanities: before any transcription can begin, the symbol inventory of the underlying cipher alphabet must first be identified and characterized. We address this challenge through symbol spotting: given a candidate alphabet specified as a set of rendered font glyphs, the task is to determine whether and where its characters appear in an unseen handwritten document, without any labeled examples from the target script. The main difficulty lies in the domain gap between clean, digitally rendered font queries and degraded handwritten manuscript symbols. We propose a three-stage pipeline that bridges this gap without manual annotation, combining a joint SimCLR+DANN encoder for domain-invariant glyph representations with an embedding-space style-adaptation mechanism applied at retrieval time, requiring no re-training. Experiments on fourteen pages from seven encrypted manuscript collections show that our method outperforms zero-shot foundation models, including CLIP and DINOv2, by a large margin ($+0.194$ P@1 over CLIP ViT-L/14), and surpasses task-specific trained baselines by $+0.138$ P@1. We further demonstrate that the Raw-Cover metric, computed in a fully unsupervised setting, provides a meaningful script-family fingerprint that identifies the underlying alphabet of an unknown document. This capability is of direct practical relevance to palaeographers, historians, and other researchers working with undeciphered manuscripts.
Problem

Research questions and friction points this paper is trying to address.

Unsupervised Domain Adaptation
Symbol Spotting
Historical Encrypted Manuscripts
Digital Humanities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised Domain Adaptation
Symbol Spotting
SimCLR+DANN Encoder
Embedding-Space Style-Adaptation
🔎 Similar Papers
No similar papers found.
G
Giuseppe De Gregorio
Computer Vision Center (CVC), Universitat Autònoma de Barcelona, Spain
A
Alicia Fornés
Computer Vision Center (CVC), Universitat Autònoma de Barcelona, Spain
Lei Kang
Lei Kang
Computer Vision Center, Universitat Autònoma de Barcelona
Document AnalysisMedical Document AnalysisDigital HumanitiesMultimodal AIGenerative AI
B
Beáta Megyesi
Stockholm University, Sweden