Institution profile

Deepnoid

Industry researchasia · kr
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

Jun 05, 2026

This work addresses the issue of “condition collapse” in digital histopathology image generation, which arises when using pretrained vision foundation models as conditioning signals. The authors propose a Riemannian flow matching–based generative framework that models normalized patch-token features as latent variables on the unit hypersphere. Their approach introduces, for the first time, a spherical-aware bridging stochastic perturbation mechanism and an anisotropic decoder, integrated within a Diffusion Transformer (DiT) architecture. Generation is further guided by a decoding strategy based on the directional energy of the Jacobian matrix of the velocity field. Evaluated on breast and colorectal cancer datasets, the method achieves state-of-the-art performance in both reconstruction fidelity and generation diversity, significantly outperforming existing approaches.

0 citationsRead paper

SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation

Jun 02, 2026

Existing policy optimization methods based on scalar advantages face challenges in long-form vision-language generation due to coarse-grained credit assignment and limited adaptability to semantically rich images. This work proposes a segment-decomposed policy optimization approach that, for the first time, incorporates the natural segmentation structure of generated outputs into multimodal reinforcement learning. It introduces verifiable segment-level rewards and employs within-segment z-normalization to produce per-segment advantage vectors, enabling fine-grained credit assignment. Built upon the GRPO framework, the method integrates segment-level normalization, verifiable rewards, and a hybrid reward strategy, significantly enhancing performance. It consistently outperforms baselines across tasks such as DOCCI, MultiChartQA, and MMSci, with particularly pronounced gains when the number of segments increases or when segments exhibit semantic independence, and it seamlessly integrates into existing systems like Dr. GRPO.

0 citationsRead paper

$MV_{Hybrid}$: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models

Aug 01, 2025

Existing Vision Transformer (ViT)-based histopathological foundation models underperform in predicting spatial gene expression (biomarkers), primarily due to their limited capacity to capture low-frequency, subtle morphological features associated with molecular phenotypes. Method: We propose $MV_{Hybrid}$, the first hybrid backbone integrating a state-space model (SSM) — initialized with negative real eigenvalues — into a Vision Transformer architecture to enhance low-frequency signal modeling. Leveraging the DINOv2 self-supervised learning framework, we systematically compare six backbone variants and adopt a leave-one-study-out external validation strategy to rigorously assess generalizability. Results: On cross-study spatial gene expression prediction, $MV_{Hybrid}$ achieves a 57% higher correlation and 43% lower performance degradation compared to the best-performing ViT baseline. Moreover, it consistently outperforms existing models across diverse downstream tasks—including classification, retrieval, and survival prediction—demonstrating superior robustness and transferability.

0 citationsRead paper

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability

Apr 10, 2025

Existing radiology multimodal models face three key challenges: weak modeling of low-resolution images, insufficient exploitation of radiology report semantics, and clinically uninterpretable cross-modal attention. This paper proposes a novel vision-language interpretable alignment framework. First, it introduces a local patch–text embedding similarity-based cross-attention mechanism—the first of its kind. Second, it designs a multi-positive contrastive learning strategy to enhance fine-grained semantic modeling of radiology reports. Third, it generates pixel-level cross-modal similarity maps that provide clinically interpretable alignment evidence and enable open-vocabulary semantic segmentation. The method integrates large language model semantic distillation with trainable Transformer layers. On public chest X-ray benchmarks, it achieves state-of-the-art zero-shot classification, localization, and segmentation performance. Extensive experiments demonstrate significant improvements in both clinical interpretability and generalization capability.

0 citationsRead paper
Recent publications

Latest Papers

STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation

Jun 05, 2026

This work addresses the issue of “condition collapse” in digital histopathology image generation, which arises when using pretrained vision foundation models as conditioning signals. The authors propose a Riemannian flow matching–based generative framework that models normalized patch-token features as latent variables on the unit hypersphere. Their approach introduces, for the first time, a spherical-aware bridging stochastic perturbation mechanism and an anisotropic decoder, integrated within a Diffusion Transformer (DiT) architecture. Generation is further guided by a decoding strategy based on the directional energy of the Jacobian matrix of the velocity field. Evaluated on breast and colorectal cancer datasets, the method achieves state-of-the-art performance in both reconstruction fidelity and generation diversity, significantly outperforming existing approaches.

0 citationsRead paper

SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation

Jun 02, 2026

Existing policy optimization methods based on scalar advantages face challenges in long-form vision-language generation due to coarse-grained credit assignment and limited adaptability to semantically rich images. This work proposes a segment-decomposed policy optimization approach that, for the first time, incorporates the natural segmentation structure of generated outputs into multimodal reinforcement learning. It introduces verifiable segment-level rewards and employs within-segment z-normalization to produce per-segment advantage vectors, enabling fine-grained credit assignment. Built upon the GRPO framework, the method integrates segment-level normalization, verifiable rewards, and a hybrid reward strategy, significantly enhancing performance. It consistently outperforms baselines across tasks such as DOCCI, MultiChartQA, and MMSci, with particularly pronounced gains when the number of segments increases or when segments exhibit semantic independence, and it seamlessly integrates into existing systems like Dr. GRPO.

0 citationsRead paper

$MV_{Hybrid}$: Improving Spatial Transcriptomics Prediction with Hybrid State Space-Vision Transformer Backbone in Pathology Vision Foundation Models

Aug 01, 2025

Existing Vision Transformer (ViT)-based histopathological foundation models underperform in predicting spatial gene expression (biomarkers), primarily due to their limited capacity to capture low-frequency, subtle morphological features associated with molecular phenotypes. Method: We propose $MV_{Hybrid}$, the first hybrid backbone integrating a state-space model (SSM) — initialized with negative real eigenvalues — into a Vision Transformer architecture to enhance low-frequency signal modeling. Leveraging the DINOv2 self-supervised learning framework, we systematically compare six backbone variants and adopt a leave-one-study-out external validation strategy to rigorously assess generalizability. Results: On cross-study spatial gene expression prediction, $MV_{Hybrid}$ achieves a 57% higher correlation and 43% lower performance degradation compared to the best-performing ViT baseline. Moreover, it consistently outperforms existing models across diverse downstream tasks—including classification, retrieval, and survival prediction—demonstrating superior robustness and transferability.

0 citationsRead paper

RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability

Apr 10, 2025

Existing radiology multimodal models face three key challenges: weak modeling of low-resolution images, insufficient exploitation of radiology report semantics, and clinically uninterpretable cross-modal attention. This paper proposes a novel vision-language interpretable alignment framework. First, it introduces a local patch–text embedding similarity-based cross-attention mechanism—the first of its kind. Second, it designs a multi-positive contrastive learning strategy to enhance fine-grained semantic modeling of radiology reports. Third, it generates pixel-level cross-modal similarity maps that provide clinically interpretable alignment evidence and enable open-vocabulary semantic segmentation. The method integrates large language model semantic distillation with trainable Transformer layers. On public chest X-ray benchmarks, it achieves state-of-the-art zero-shot classification, localization, and segmentation performance. Extensive experiments demonstrate significant improvements in both clinical interpretability and generalization capability.

0 citationsRead paper