A deep dictionary network-based foundation model for ultra-low-dose CT denoising
为解决超低剂量CT图像噪声问题,提出基于深度字典网络的统一多器官去噪基础模型,通过预训练和微调实现跨区域去噪。
为解决超低剂量CT图像噪声问题,提出基于深度字典网络的统一多器官去噪基础模型,通过预训练和微调实现跨区域去噪。
论文针对医学评估中逻辑一致性问题,通过建立LogiMed-RoB基准测试来评估大型语言模型的逻辑推理能力,揭示了现有模型在深层逻辑推理上的不足。
本文提出一种基于梯度但依赖于尖峰时序的方法,通过状态分离和梯度隧道算法解决神经微电路中的时间信用分配问题。
为解决血管介入手术中造影剂风险问题,提出VeCAS框架,通过无造影剂X射线图像生成血管造影图,提高手术效率和安全性。
This study addresses the challenges of weak temporal reasoning, detail loss, and poor transferability in long-duration surgical video understanding by proposing a perception-reasoning decoupled agent framework. The architecture employs a text orchestrator to plan evidence collection alongside frozen visual sub-agents for tool execution, integrated with a gradient-free heuristic skill distillation mechanism that adaptively evolves a reusable external skill library from low-scoring trajectories. Requiring only approximately one hundred annotated samples for skill retrieval optimization, this approach comprehensively outperforms existing vision-language models and video agents on both proprietary neurosurgical and public benchmarks. Consequently, the proposed method significantly enhances out-of-domain generalization capabilities and temporal reasoning accuracy within long-video contexts, offering a robust solution for complex surgical analysis tasks.
为解决超低剂量CT图像噪声问题,提出基于深度字典网络的统一多器官去噪基础模型,通过预训练和微调实现跨区域去噪。
论文针对医学评估中逻辑一致性问题,通过建立LogiMed-RoB基准测试来评估大型语言模型的逻辑推理能力,揭示了现有模型在深层逻辑推理上的不足。
本文提出一种基于梯度但依赖于尖峰时序的方法,通过状态分离和梯度隧道算法解决神经微电路中的时间信用分配问题。
为解决血管介入手术中造影剂风险问题,提出VeCAS框架,通过无造影剂X射线图像生成血管造影图,提高手术效率和安全性。
This study addresses the challenges of weak temporal reasoning, detail loss, and poor transferability in long-duration surgical video understanding by proposing a perception-reasoning decoupled agent framework. The architecture employs a text orchestrator to plan evidence collection alongside frozen visual sub-agents for tool execution, integrated with a gradient-free heuristic skill distillation mechanism that adaptively evolves a reusable external skill library from low-scoring trajectories. Requiring only approximately one hundred annotated samples for skill retrieval optimization, this approach comprehensively outperforms existing vision-language models and video agents on both proprietary neurosurgical and public benchmarks. Consequently, the proposed method significantly enhances out-of-domain generalization capabilities and temporal reasoning accuracy within long-video contexts, offering a robust solution for complex surgical analysis tasks.