DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning
为解决多模态牙科诊断准确性问题,提出DentAgent框架,通过协调五个专业代理处理不同数据类型,并利用共享证据状态提高诊断的可追溯性和准确性。
为解决多模态牙科诊断准确性问题,提出DentAgent框架,通过协调五个专业代理处理不同数据类型,并利用共享证据状态提高诊断的可追溯性和准确性。
This study addresses the limitations of conventional cone-beam computed tomography (CBCT)—notably high radiation exposure and cost—and the inability of single panoramic X-rays to yield geometrically consistent, high-fidelity 3D dental reconstructions. To overcome these challenges, the authors propose HiCT, a two-stage framework that first leverages a video diffusion model to synthesize geometrically consistent multi-view projections from a single panoramic X-ray, followed by high-fidelity CBCT reconstruction via a ray-based dynamic attention network integrated with an X-ray sampling strategy. This work pioneers the integration of video diffusion models with ray-wise dynamic attention mechanisms and introduces XCT, a large-scale paired dataset enabling robust training and validation. Experimental results demonstrate state-of-the-art performance on clinically relevant metrics, achieving accurate, geometrically consistent CBCT reconstructions with strong potential for clinical translation.
Surgical video understanding has been hindered by limited data scale, narrow procedural diversity, inconsistent evaluation protocols, and non-reproducible training pipelines. To address these challenges, this work proposes SurgRec—a scalable and reproducible self-supervised pretraining framework for surgical videos, featuring two variants: SurgRec-MAE and SurgRec-JEPA. The study introduces the first large-scale, multi-source surgical video corpus encompassing diverse procedures, integrated with a balanced sampling strategy and a unified downstream evaluation benchmark. This approach substantially enhances model generalization across tasks. Evaluated on 16 downstream datasets, SurgRec consistently outperforms existing self-supervised and vision-language methods, demonstrating particularly robust performance in fine-grained temporal recognition tasks.
This work addresses the limitation of existing dental vision-language models in effectively leveraging the native 3D geometric information from intraoral scans (IOS), which hinders unified multi-disease diagnosis. To overcome this, we propose IOSVLM, an end-to-end 3D vision-language model that represents IOS as point clouds and integrates a 3D encoder, a projector, and a large language model to enable unified diagnosis and generative visual question answering grounded in 3D geometry. To bridge the distribution gap between colorless IOS data and color-dependent 3D pretraining, we design a geometry-to-color proxy mechanism and adopt a two-stage curriculum learning strategy to enhance robustness. We also introduce IOSVQA, a large-scale, multi-source VQA dataset for IOS-based diagnosis. Experiments show that IOSVLM significantly outperforms strong baselines, achieving a 9.58% gain in macro accuracy and a 1.46% improvement in macro F1, validating the efficacy of directly modeling 3D geometry.
Clinical decision-making is often inefficient and prone to missed diagnoses due to challenges in fusing heterogeneous multimodal medical data—such as text, 2D/3D imaging, and video. Existing medical vision-language models (VLMs) suffer from architectural opacity, scarcity of high-quality annotations, and poor scalability across modalities. To address these limitations, we propose the first transparent, unified, full-modality medical VLM framework. Our approach introduces a medical-aware token compression mechanism and a progressive multi-scale patch encoder, enabling synergistic learning across 2D → 3D → video modalities. We employ end-to-end alignment training with efficient token reduction. Evaluated on 30 cross-modal medical benchmarks, our method achieves state-of-the-art performance. Models ranging from 7B to 32B parameters require only 4K–40K GPU-hours for training—matching or surpassing closed-source systems in accuracy while significantly enhancing clinical interpretability and deployment flexibility.
为解决多模态牙科诊断准确性问题,提出DentAgent框架,通过协调五个专业代理处理不同数据类型,并利用共享证据状态提高诊断的可追溯性和准确性。
This study addresses the limitations of conventional cone-beam computed tomography (CBCT)—notably high radiation exposure and cost—and the inability of single panoramic X-rays to yield geometrically consistent, high-fidelity 3D dental reconstructions. To overcome these challenges, the authors propose HiCT, a two-stage framework that first leverages a video diffusion model to synthesize geometrically consistent multi-view projections from a single panoramic X-ray, followed by high-fidelity CBCT reconstruction via a ray-based dynamic attention network integrated with an X-ray sampling strategy. This work pioneers the integration of video diffusion models with ray-wise dynamic attention mechanisms and introduces XCT, a large-scale paired dataset enabling robust training and validation. Experimental results demonstrate state-of-the-art performance on clinically relevant metrics, achieving accurate, geometrically consistent CBCT reconstructions with strong potential for clinical translation.
Surgical video understanding has been hindered by limited data scale, narrow procedural diversity, inconsistent evaluation protocols, and non-reproducible training pipelines. To address these challenges, this work proposes SurgRec—a scalable and reproducible self-supervised pretraining framework for surgical videos, featuring two variants: SurgRec-MAE and SurgRec-JEPA. The study introduces the first large-scale, multi-source surgical video corpus encompassing diverse procedures, integrated with a balanced sampling strategy and a unified downstream evaluation benchmark. This approach substantially enhances model generalization across tasks. Evaluated on 16 downstream datasets, SurgRec consistently outperforms existing self-supervised and vision-language methods, demonstrating particularly robust performance in fine-grained temporal recognition tasks.
This work addresses the limitation of existing dental vision-language models in effectively leveraging the native 3D geometric information from intraoral scans (IOS), which hinders unified multi-disease diagnosis. To overcome this, we propose IOSVLM, an end-to-end 3D vision-language model that represents IOS as point clouds and integrates a 3D encoder, a projector, and a large language model to enable unified diagnosis and generative visual question answering grounded in 3D geometry. To bridge the distribution gap between colorless IOS data and color-dependent 3D pretraining, we design a geometry-to-color proxy mechanism and adopt a two-stage curriculum learning strategy to enhance robustness. We also introduce IOSVQA, a large-scale, multi-source VQA dataset for IOS-based diagnosis. Experiments show that IOSVLM significantly outperforms strong baselines, achieving a 9.58% gain in macro accuracy and a 1.46% improvement in macro F1, validating the efficacy of directly modeling 3D geometry.
Clinical decision-making is often inefficient and prone to missed diagnoses due to challenges in fusing heterogeneous multimodal medical data—such as text, 2D/3D imaging, and video. Existing medical vision-language models (VLMs) suffer from architectural opacity, scarcity of high-quality annotations, and poor scalability across modalities. To address these limitations, we propose the first transparent, unified, full-modality medical VLM framework. Our approach introduces a medical-aware token compression mechanism and a progressive multi-scale patch encoder, enabling synergistic learning across 2D → 3D → video modalities. We employ end-to-end alignment training with efficient token reduction. Evaluated on 30 cross-modal medical benchmarks, our method achieves state-of-the-art performance. Models ranging from 7B to 32B parameters require only 4K–40K GPU-hours for training—matching or surpassing closed-source systems in accuracy while significantly enhancing clinical interpretability and deployment flexibility.