Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of error accumulation, high computational cost, and strong data dependency in multi-oriented scene text recognition by proposing RISTER, an end-to-end rotation-invariant network. The encoder integrates rotation-equivariant convolutions with self-attention to jointly capture local and global features, while the decoder is designed to be fully rotation-invariant based on a novel theoretical proof that cross-attention inherently possesses rotation invariance. Evaluated on standard and multi-oriented text benchmarks, RISTER achieves state-of-the-art performance, surpassing the second-best model by 4.0% in accuracy on general multi-oriented datasets without incurring additional inference overhead.
📝 Abstract
Multi-oriented text is ubiquitous in real-world scenes and remains a major challenge for scene text recognition (STR). Existing rotation-aware methods explicitly estimate text orientation. However, due to the lack of theoretical guarantees, they are prone to error accumulation, increased computational cost, and strong reliance on data. In this work, we incorporate rotation invariance into the STR framework to address these limitations. Specifically, we adopt an encoder-decoder architecture, embedding rotation equivariance in the encoder and rotation invariance in the decoder to construct a fully rotation-invariant network. On the decoder side, we first identify and prove the rotation-invariant property of the cross-attention mechanism and use it to formulate a rotation-invariant text decoder that maps visual features to output text in a rotation-invariant manner. On the encoder side, we propose a rotation-equivariant local-global extraction network that integrates deep equivariant convolutions with self-attention, enabling rotation-equivariant feature extraction while modeling inter-character dependencies and preserving fine-grained visual details. By integrating the encoder and decoder, we obtain an end-to-end Rotation-Invariant Scene Text Recognition network (RISTER). RISTER provides rotation invariance with theoretical guarantees, enhancing robustness on multi-oriented samples without introducing additional inference computation or relying on data-driven orientation correction. Experiments show that RISTER achieves state-of-the-art performance on both standard and multi-oriented benchmarks, surpassing the second-best model by 4.0 percent in accuracy on the general multi-oriented dataset.
Problem

Research questions and friction points this paper is trying to address.

scene text recognition
rotation invariance
multi-oriented text
rotation-aware methods
text orientation
Innovation

Methods, ideas, or system contributions that make the work stand out.

rotation invariance
scene text recognition
equivariant convolution
cross-attention
multi-oriented text
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhibin Ma
Shenzhen Campus of Sun Yat-sen University, China; Shenzhen Key Laboratory of Adversarial Artificial Intelligence, China
P
Pengwen Dai
Shenzhen Campus of Sun Yat-sen University, China; Shenzhen Key Laboratory of Adversarial Artificial Intelligence, China
Yi Liu
Yi Liu
Baidu Inc.
CVLLMVLM
Xugong Qin
Xugong Qin
Nanjing University of Science and Technology
Computer VisionDocument AnalysisMedia Forensics
Chenyun Yu
Chenyun Yu
phd, Department of Computer Science, City University of Hong Kong
Data science and managementquery optimizationdata mininginformation security
Xiaochun Cao
Xiaochun Cao
Sun Yat-sen University
Computer VisionArtificial IntelligenceMultimediaMachine Learning