Institution profile

Angelalign Technology Inc.

Industry researchasia · cn
Official website
Research library11linked papers
Opportunities0open roles
Selected work

Representative Papers

HICT: High-precision 3D CBCT reconstruction from a single X-ray

Apr 01, 2026

This study addresses the limitations of conventional cone-beam computed tomography (CBCT)—notably high radiation exposure and cost—and the inability of single panoramic X-rays to yield geometrically consistent, high-fidelity 3D dental reconstructions. To overcome these challenges, the authors propose HiCT, a two-stage framework that first leverages a video diffusion model to synthesize geometrically consistent multi-view projections from a single panoramic X-ray, followed by high-fidelity CBCT reconstruction via a ray-based dynamic attention network integrated with an X-ray sampling strategy. This work pioneers the integration of video diffusion models with ray-wise dynamic attention mechanisms and introduces XCT, a large-scale paired dataset enabling robust training and validation. Experimental results demonstrate state-of-the-art performance on clinically relevant metrics, achieving accurate, geometrically consistent CBCT reconstructions with strong potential for clinical translation.

0 citationsRead paper

Scaling Video Pretraining for Surgical Foundation Models

Mar 31, 2026

Surgical video understanding has been hindered by limited data scale, narrow procedural diversity, inconsistent evaluation protocols, and non-reproducible training pipelines. To address these challenges, this work proposes SurgRec—a scalable and reproducible self-supervised pretraining framework for surgical videos, featuring two variants: SurgRec-MAE and SurgRec-JEPA. The study introduces the first large-scale, multi-source surgical video corpus encompassing diverse procedures, integrated with a balanced sampling strategy and a unified downstream evaluation benchmark. This approach substantially enhances model generalization across tasks. Evaluated on 16 downstream datasets, SurgRec consistently outperforms existing self-supervised and vision-language methods, demonstrating particularly robust performance in fine-grained temporal recognition tasks.

0 citationsRead paper

IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans

Mar 17, 2026

This work addresses the limitation of existing dental vision-language models in effectively leveraging the native 3D geometric information from intraoral scans (IOS), which hinders unified multi-disease diagnosis. To overcome this, we propose IOSVLM, an end-to-end 3D vision-language model that represents IOS as point clouds and integrates a 3D encoder, a projector, and a large language model to enable unified diagnosis and generative visual question answering grounded in 3D geometry. To bridge the distribution gap between colorless IOS data and color-dependent 3D pretraining, we design a geometry-to-color proxy mechanism and adopt a two-stage curriculum learning strategy to enhance robustness. We also introduce IOSVQA, a large-scale, multi-source VQA dataset for IOS-based diagnosis. Experiments show that IOSVLM significantly outperforms strong baselines, achieving a 9.58% gain in macro accuracy and a 1.46% improvement in macro F1, validating the efficacy of directly modeling 3D geometry.

0 citationsRead paper

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Oct 09, 2025

Clinical decision-making is often inefficient and prone to missed diagnoses due to challenges in fusing heterogeneous multimodal medical data—such as text, 2D/3D imaging, and video. Existing medical vision-language models (VLMs) suffer from architectural opacity, scarcity of high-quality annotations, and poor scalability across modalities. To address these limitations, we propose the first transparent, unified, full-modality medical VLM framework. Our approach introduces a medical-aware token compression mechanism and a progressive multi-scale patch encoder, enabling synergistic learning across 2D → 3D → video modalities. We employ end-to-end alignment training with efficient token reduction. Evaluated on 30 cross-modal medical benchmarks, our method achieves state-of-the-art performance. Models ranging from 7B to 32B parameters require only 4K–40K GPU-hours for training—matching or surpassing closed-source systems in accuracy while significantly enhancing clinical interpretability and deployment flexibility.

0 citationsRead paper
Recent publications

Latest Papers

HICT: High-precision 3D CBCT reconstruction from a single X-ray

Apr 01, 2026

This study addresses the limitations of conventional cone-beam computed tomography (CBCT)—notably high radiation exposure and cost—and the inability of single panoramic X-rays to yield geometrically consistent, high-fidelity 3D dental reconstructions. To overcome these challenges, the authors propose HiCT, a two-stage framework that first leverages a video diffusion model to synthesize geometrically consistent multi-view projections from a single panoramic X-ray, followed by high-fidelity CBCT reconstruction via a ray-based dynamic attention network integrated with an X-ray sampling strategy. This work pioneers the integration of video diffusion models with ray-wise dynamic attention mechanisms and introduces XCT, a large-scale paired dataset enabling robust training and validation. Experimental results demonstrate state-of-the-art performance on clinically relevant metrics, achieving accurate, geometrically consistent CBCT reconstructions with strong potential for clinical translation.

0 citationsRead paper

Scaling Video Pretraining for Surgical Foundation Models

Mar 31, 2026

Surgical video understanding has been hindered by limited data scale, narrow procedural diversity, inconsistent evaluation protocols, and non-reproducible training pipelines. To address these challenges, this work proposes SurgRec—a scalable and reproducible self-supervised pretraining framework for surgical videos, featuring two variants: SurgRec-MAE and SurgRec-JEPA. The study introduces the first large-scale, multi-source surgical video corpus encompassing diverse procedures, integrated with a balanced sampling strategy and a unified downstream evaluation benchmark. This approach substantially enhances model generalization across tasks. Evaluated on 16 downstream datasets, SurgRec consistently outperforms existing self-supervised and vision-language methods, demonstrating particularly robust performance in fine-grained temporal recognition tasks.

0 citationsRead paper

IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans

Mar 17, 2026

This work addresses the limitation of existing dental vision-language models in effectively leveraging the native 3D geometric information from intraoral scans (IOS), which hinders unified multi-disease diagnosis. To overcome this, we propose IOSVLM, an end-to-end 3D vision-language model that represents IOS as point clouds and integrates a 3D encoder, a projector, and a large language model to enable unified diagnosis and generative visual question answering grounded in 3D geometry. To bridge the distribution gap between colorless IOS data and color-dependent 3D pretraining, we design a geometry-to-color proxy mechanism and adopt a two-stage curriculum learning strategy to enhance robustness. We also introduce IOSVQA, a large-scale, multi-source VQA dataset for IOS-based diagnosis. Experiments show that IOSVLM significantly outperforms strong baselines, achieving a 9.58% gain in macro accuracy and a 1.46% improvement in macro F1, validating the efficacy of directly modeling 3D geometry.

0 citationsRead paper

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

Oct 09, 2025

Clinical decision-making is often inefficient and prone to missed diagnoses due to challenges in fusing heterogeneous multimodal medical data—such as text, 2D/3D imaging, and video. Existing medical vision-language models (VLMs) suffer from architectural opacity, scarcity of high-quality annotations, and poor scalability across modalities. To address these limitations, we propose the first transparent, unified, full-modality medical VLM framework. Our approach introduces a medical-aware token compression mechanism and a progressive multi-scale patch encoder, enabling synergistic learning across 2D → 3D → video modalities. We employ end-to-end alignment training with efficient token reduction. Evaluated on 30 cross-modal medical benchmarks, our method achieves state-of-the-art performance. Models ranging from 7B to 32B parameters require only 4K–40K GPU-hours for training—matching or surpassing closed-source systems in accuracy while significantly enhancing clinical interpretability and deployment flexibility.

0 citationsRead paper