A Hybrid Vision Transformer Approach for Mathematical Expression Recognition

πŸ“… 2022-11-30
πŸ›οΈ International Conference on Digital Image Computing: Techniques and Applications
πŸ“ˆ Citations: 2
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work proposes a hybrid Vision Transformer-based sequence-to-sequence approach to address the challenges in mathematical expression recognition arising from the two-dimensional layout and scale variations of symbols. The encoder integrates 2D positional encoding to effectively capture spatial dependencies among symbols and innovatively employs the ViT’s [CLS] token as the initial embedding for the decoder. To mitigate issues of under- or over-parsing during decoding, a coverage attention mechanism is introduced. Evaluated on the IM2LATEX-100K dataset, the proposed method achieves a BLEU score of 89.94, outperforming current state-of-the-art approaches and significantly improving the accuracy of formula recognition.

Technology Category

Application Category

πŸ“ Abstract
Mathematical expression recognition is one of the important processes in scientific documents analysis. Despite the importance of this task, solving mathematical expression recognition is still very challenging. One of the reasons for the difficulty of math recognition compared to normal text recognition is that math formula usually has 2-D spatial structure relationship [1] instead of 1-D ones from normal text data. The spatial structure relationship of math formula is presented by many math symbols such as superscript, subscript, fraction symbol, etc. The traditional approach usually solves this problem in two stages. First, the character segmentation stage is used to segment each character in math formula and then classify it based on the given vocabulary. Second, the structural analysis stage is used to identify the spatial relationships between all characters of the math formula.
Problem

Research questions and friction points this paper is trying to address.

mathematical expression recognition
document analysis
two-dimensional structure
symbol size variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Vision Transformer
2D positional encoding
coverage attention
mathematical expression recognition
[CLS] token
Anh Duy Le
Anh Duy Le
Viettel AI
Multimodal LearningGenerative Modeling
V
Van Linh Pham
Viettel Cyberspace Center, Viettel Group, Vietnam. Lot D26 Cau Giay New Urban Area, Yen Hoa Ward, Cau Giay District, Hanoi, Vietnam.
V
Vinh Loi Ly
Viettel Cyberspace Center, Viettel Group, Vietnam. Lot D26 Cau Giay New Urban Area, Yen Hoa Ward, Cau Giay District, Hanoi, Vietnam.
N
Nam Quan Nguyen
Viettel Cyberspace Center, Viettel Group, Vietnam. Lot D26 Cau Giay New Urban Area, Yen Hoa Ward, Cau Giay District, Hanoi, Vietnam.
H
Huu Thang Nguyen
Faculty of Computer Science&Engineering, Ho Chi Minh City-University of Technology (HCMUT), 268 Ly Thuong Kiet Street, District 10, Ho Chi Minh City, Vietnam.; Vietnam National University Ho Chi Minh City, Linh Trung Ward, Thu Duc District, Ho Chi Minh City, Vietnam.
T
Tuan Anh Tran
Faculty of Computer Science&Engineering, Ho Chi Minh City-University of Technology (HCMUT), 268 Ly Thuong Kiet Street, District 10, Ho Chi Minh City, Vietnam.; Vietnam National University Ho Chi Minh City, Linh Trung Ward, Thu Duc District, Ho Chi Minh City, Vietnam.