Bag of Bags: Adaptive Visual Vocabularies for Genizah Join Image Retrieval

📅 2026-04-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of reconnecting fragmented manuscripts from the Cairo Genizah by proposing a "Bag-of-Bags" (BoB) representation. Instead of relying on a global visual vocabulary, BoB constructs an adaptive local visual codebook for each image. Binary patch features are extracted via a sparse convolutional autoencoder, and set representations are formed by combining connected component analysis with image-level k-means clustering. Fragment matching is performed based on distances between these sets. The authors further introduce BoB-OT, a variant that incorporates cluster-size weighting and provides theoretical approximation guarantees with respect to full-component optimal transport discrepancy. Evaluated on real-world data, their two-stage retrieval pipeline—comprising an initial Bag-of-Words filtering followed by BoB-OT reranking—achieves a Hit@1 of 0.78 and MRR of 0.84, outperforming the strongest BoW baseline by 6.1% in top-1 accuracy while maintaining computational efficiency.

Technology Category

Application Category

📝 Abstract
A join is a set of manuscript fragments identified as originally emanating from the same manuscript. We study manuscript join retrieval: Given a query image of a fragment, retrieve other fragments originating from the same physical manuscript. We propose Bag of Bags (BoB), an image-level representation that replaces the global-level visual codebook of classical Bag of Words (BoW) with a fragment-specific vocabulary of local visual words. Our pipeline trains a sparse convolutional autoencoder on binarized fragment patches, encodes connected components from each page, clusters the resulting embeddings with per image $k$-means, and compares images using set to set distances between their local vocabularies. Evaluated on fragments from the Cairo Genizah, the best BoB variant (viz.\@ Chamfer) achieves Hit@1 of 0.78 and MRR of 0.84, compared to 0.74 and 0.80, respectively, for the strongest BoW baseline (BoW-RawPatches-$χ^2$), a 6.1\% relative improvement in top-1 accuracy. We furthermore study a mass-weighted BoB-OT variant that incorporates cluster population into prototype matching and present a formal approximation guarantee bounding its deviation from full component-level optimal transport. A two-stage pipeline using a BoW shortlist followed by BoB-OT reranking provides a practical compromise between retrieval strength and computational cost, supporting applicability to larger manuscript collections.
Problem

Research questions and friction points this paper is trying to address.

manuscript join retrieval
fragment matching
visual vocabulary
image retrieval
Cairo Genizah
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bag of Bags
fragment-specific vocabulary
sparse convolutional autoencoder
set-to-set distance
optimal transport
🔎 Similar Papers
No similar papers found.
Sharva Gogawale
Sharva Gogawale
Tel Aviv University
Computer VisionAI for Social GoodComputational HumanitiesHuman-Computer Interaction
G
Gal Grudka
School of Computer Science and AI, Tel Aviv University, Ramat Aviv, Israel
D
Daria Vasyutinsky-Shapira
School of Computer Science and AI, Tel Aviv University, Ramat Aviv, Israel
O
Omer Ventura
School of Computer Science and AI, Tel Aviv University, Ramat Aviv, Israel
B
Berat Kurar-Barakat
School of Computer Science and AI, Tel Aviv University, Ramat Aviv, Israel
N
Nachum Dershowitz
School of Computer Science and AI, Tel Aviv University, Ramat Aviv, Israel