Systematic Literature Review of Machine Learning Models and Applications for Text Recognition
该文通过系统性文献回顾,评估了近十年OCR技术的发展,分析了97篇相关研究,探讨了解决多语言处理和复杂数据格式的方法及挑战。
该文通过系统性文献回顾,评估了近十年OCR技术的发展,分析了97篇相关研究,探讨了解决多语言处理和复杂数据格式的方法及挑战。
本文提出一种基于监督学习的框架,通过RTL级设计参数直接预测MBIST面积和测试时间,提高了预测准确性和效率。
This study addresses the high privacy risks inherent in complex tropical traffic scenarios in Kuala Lumpur by proposing an automated anonymization framework. By integrating Grounding DINO with a novel spatial vehicle ROI constraint mechanism, temporal persistence, and automated quality inspection, the method effectively suppresses environmental false positives and enhances occluded object recognition while preserving scene context. Experimental evaluation on a 1,266-frame test set demonstrates an anonymization success rate of approximately 95%, significantly mitigating privacy leakage risks. Consequently, this approach provides a robust solution for the efficient de-identification of urban traffic datasets in tropical environments, balancing data utility with stringent privacy protection requirements.
This work addresses the lack of verifiable guarantees—such as convergence, interpretability, and bounded interaction rounds—in multi-agent large language model reasoning. The authors model collective reasoning as a nonlinear dynamical system over a communication graph and, for the first time, apply Koopman operator theory to construct a linear representation from interaction trajectories. Spectral analysis of this representation yields three machine-verifiable certificates: convergence deadlines, identification of cohesive cliques with interpretable validity, and an auditable basis for compressed messages. Experiments demonstrate that convergence rounds are predicted accurately in 96% of configurations (log-scale correlation of 0.93), attributions are exact, and decision-relevant information is preserved with 99.7% fidelity using only 8 out of 32 spectral coordinates. Certificates trained on 15 debates remain fully valid across 60 leave-one-out tests and are computable within minutes on a CPU.
This work addresses the challenge of barren plateaus in variational quantum algorithms for medical image classification, which lead to vanishing gradients and hinder training. The authors propose a novel approach that integrates large language model (LLM)-guided single-query initialization with AdaInit, CUDA-Q GPU-accelerated quantum simulation, and prompt engineering to generate high-quality initial parameters without iterative optimization. Applied to binary classification of mammogram images, this method avoids barren plateaus effectively, yielding a 14.6× increase in gradient variance and a 160× acceleration in convergence time (1.1 seconds versus 176 seconds) compared to random initialization, while maintaining a classification accuracy of 61.4%. These results demonstrate a significant improvement in the trainability and efficiency of hybrid quantum-classical models.
该文通过系统性文献回顾,评估了近十年OCR技术的发展,分析了97篇相关研究,探讨了解决多语言处理和复杂数据格式的方法及挑战。
本文提出一种基于监督学习的框架,通过RTL级设计参数直接预测MBIST面积和测试时间,提高了预测准确性和效率。
This study addresses the high privacy risks inherent in complex tropical traffic scenarios in Kuala Lumpur by proposing an automated anonymization framework. By integrating Grounding DINO with a novel spatial vehicle ROI constraint mechanism, temporal persistence, and automated quality inspection, the method effectively suppresses environmental false positives and enhances occluded object recognition while preserving scene context. Experimental evaluation on a 1,266-frame test set demonstrates an anonymization success rate of approximately 95%, significantly mitigating privacy leakage risks. Consequently, this approach provides a robust solution for the efficient de-identification of urban traffic datasets in tropical environments, balancing data utility with stringent privacy protection requirements.
This work addresses the lack of verifiable guarantees—such as convergence, interpretability, and bounded interaction rounds—in multi-agent large language model reasoning. The authors model collective reasoning as a nonlinear dynamical system over a communication graph and, for the first time, apply Koopman operator theory to construct a linear representation from interaction trajectories. Spectral analysis of this representation yields three machine-verifiable certificates: convergence deadlines, identification of cohesive cliques with interpretable validity, and an auditable basis for compressed messages. Experiments demonstrate that convergence rounds are predicted accurately in 96% of configurations (log-scale correlation of 0.93), attributions are exact, and decision-relevant information is preserved with 99.7% fidelity using only 8 out of 32 spectral coordinates. Certificates trained on 15 debates remain fully valid across 60 leave-one-out tests and are computable within minutes on a CPU.
This work addresses the challenge of barren plateaus in variational quantum algorithms for medical image classification, which lead to vanishing gradients and hinder training. The authors propose a novel approach that integrates large language model (LLM)-guided single-query initialization with AdaInit, CUDA-Q GPU-accelerated quantum simulation, and prompt engineering to generate high-quality initial parameters without iterative optimization. Applied to binary classification of mammogram images, this method avoids barren plateaus effectively, yielding a 14.6× increase in gradient variance and a 160× acceleration in convergence time (1.1 seconds versus 176 seconds) compared to random initialization, while maintaining a classification accuracy of 61.4%. These results demonstrate a significant improvement in the trainability and efficiency of hybrid quantum-classical models.