ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts
为解决手写文本识别资源分散、标注格式不一的问题,提出ExpertHTR框架,通过多任务学习和稀疏混合专家模型提升识别效果。
为解决手写文本识别资源分散、标注格式不一的问题,提出ExpertHTR框架,通过多任务学习和稀疏混合专家模型提升识别效果。
本文提出一种验证器引导的可解释推理框架,通过黄金锚定QLoRA、任务感知混合专家系统和组相对RLVR方法,提高大语言模型在教育问答中的解释性和准确性。
为解决0.064T下儿科脑MRI分割难题,提出AURA方法,利用非对称监督策略处理高场和低场标注差异。
This study addresses the ill-posed inverse problem in sparse-view CT reconstruction and the poor convergence of existing deep learning methods by proposing a compact deep unfolding framework inspired by second-order optimization. By constructing a structured Hessian proxy with a conjugate gradient solver and designing a global-local regularization module that integrates convolutional features with Nyström attention, the method effectively models image priors. Experiments on AAPM and DeepLesion datasets demonstrate stable convergence, significant noise power reduction, and enhanced visual fidelity. Achieving superior quantitative metrics compared to state-of-the-art approaches, this work provides an efficient and reliable solution for sparse-view CT reconstruction.
This study addresses the challenges of autonomous agent execution and test case generation in natural language-driven web automation testing by proposing an iterative planning agent based on large language models. The proposed method innovatively integrates an active error correction strategy with a multi-source memory mechanism, effectively synthesizing short-term action feedback and long-term experiential knowledge to enable end-to-end autonomous testing. Experimental results demonstrate that the agent achieves an accuracy of 97.4% on the MiniWoB++ benchmark and 83.8% across a 350-task suite. These outcomes significantly outperform existing baselines such as WALT, validating the approach’s effectiveness in complex web interaction scenarios.
为解决手写文本识别资源分散、标注格式不一的问题,提出ExpertHTR框架,通过多任务学习和稀疏混合专家模型提升识别效果。
本文提出一种验证器引导的可解释推理框架,通过黄金锚定QLoRA、任务感知混合专家系统和组相对RLVR方法,提高大语言模型在教育问答中的解释性和准确性。
为解决0.064T下儿科脑MRI分割难题,提出AURA方法,利用非对称监督策略处理高场和低场标注差异。
This study addresses the ill-posed inverse problem in sparse-view CT reconstruction and the poor convergence of existing deep learning methods by proposing a compact deep unfolding framework inspired by second-order optimization. By constructing a structured Hessian proxy with a conjugate gradient solver and designing a global-local regularization module that integrates convolutional features with Nyström attention, the method effectively models image priors. Experiments on AAPM and DeepLesion datasets demonstrate stable convergence, significant noise power reduction, and enhanced visual fidelity. Achieving superior quantitative metrics compared to state-of-the-art approaches, this work provides an efficient and reliable solution for sparse-view CT reconstruction.
This study addresses the challenges of autonomous agent execution and test case generation in natural language-driven web automation testing by proposing an iterative planning agent based on large language models. The proposed method innovatively integrates an active error correction strategy with a multi-source memory mechanism, effectively synthesizing short-term action feedback and long-term experiential knowledge to enable end-to-end autonomous testing. Experimental results demonstrate that the agent achieves an accuracy of 97.4% on the MiniWoB++ benchmark and 83.8% across a 350-task suite. These outcomes significantly outperform existing baselines such as WALT, validating the approach’s effectiveness in complex web interaction scenarios.