Institution profile

Università Campus Bio-Medico di Roma

Academic institutioneurope · it
Official website
Research library42linked papers
Opportunities0open roles
Selected work

Representative Papers

Probabilistic NDVI Forecasting from Sparse Satellite Time Series and Weather Covariates

Feb 04, 2026arXiv.org

This study addresses the challenges of sparse and irregular satellite NDVI observations caused by cloud cover and the difficulty of short-term forecasting of crop vegetation dynamics under heterogeneous climatic conditions. The authors propose a probabilistic forecasting framework that employs a deep learning architecture to separately encode historical NDVI and meteorological observations along with future exogenous covariates, fusing multimodal information for multi-step quantile prediction. A novel temporally distance-weighted quantile loss function is introduced, complemented by feature engineering that incorporates both cumulative and extreme weather metrics, effectively capturing the delayed vegetation response to meteorological drivers and temporal uncertainty. Experiments on European satellite data demonstrate that the proposed method outperforms existing statistical, deep learning, and time series baselines in both point and probabilistic forecasting metrics, with ablation studies confirming historical NDVI as the dominant predictor and meteorological covariates providing significant performance gains.

1 citationsRead paper

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

Aug 08, 2026

Existing cross-modal medical image translation methods typically rely on 2D slices or 3D local patches and require separate model training for each task, resulting in limited generalization. This work proposes a whole-volume representation learning framework based on a 3D variational autoencoder, formulating cross-modal translation as a conditional flow-matching problem in latent space. By integrating resolution-aware sampling with multi-task joint training, the approach enables a single model to support diverse modality conversions. It achieves, for the first time, unified whole-volume, multi-task translation with less than 0.15 SSIM performance drop on unseen anatomical regions. The method further supports zero-shot anatomical generalization and unsupervised cross-dataset compositional translation, matching the performance of task-specific models.

0 citationsRead paper

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

Aug 06, 2026

This study investigates whether self-supervised pre-training (SPT) can effectively enhance the diagnostic performance of Transformer models on multimodal, multivariate, and univariate medical time-series data, particularly in data-scarce clinical settings. The work proposes a general-purpose approach that requires no task-specific architectural modifications and incorporates four masking strategies to facilitate representation learning across both modalities and temporal dimensions. Experimental results across three medical time-series tasks demonstrate that SPT improves classification accuracy by 0–6 percentage points, with deeper models exhibiting more pronounced gains. Notably, this is the first study to validate the efficacy of SPT on univariate medical time-series data, demonstrating its strong generalizability, scalability, and robustness under limited-data conditions.

0 citationsRead paper

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Jun 29, 2026

This study addresses the limitations of current radiology report generation models, which rely on holistic evaluation metrics that fail to verify whether diagnoses are grounded in genuine pathological visual evidence, rendering them susceptible to spurious correlations or prior biases. To tackle this issue, the authors introduce the SHOVIR benchmark, incorporating region-level CheXpert labels on MIMIC-CXR and PadChest-GR datasets, and design image- and disease-level occlusion experiments to systematically distinguish between “direct shortcuts” and “contextual shortcuts”—two failure modes of visual dependency. Through a spatially aligned evaluation framework, the work exposes a critical blind spot in existing methods: their neglect of fine-grained regional perception. Experiments across eight state-of-the-art vision-language models reveal that fluent report generation does not necessarily indicate reliable visual grounding, as top-performing models may still exhibit shallow reliance on image evidence.

0 citationsRead paper

VegSim: A Geospatial World Model for Scenario-Conditioned Vegetation Simulation

Jun 20, 2026

Existing vegetation prediction models are constrained by fixed observational meteorological trajectories, limiting their capacity for multi-scenario response analysis. This work proposes the first geospatial world model, which integrates sparse NDVI time series, historical meteorological covariates, and static geographic context through a recurrent latent dynamics architecture. Without requiring scenario-specific supervision, the model simultaneously enables probabilistic forecasting under observed conditions and conditional simulation under user-defined meteorological forcings. Evaluated on the GreenEarthNet benchmark, it significantly outperforms both temporal and remote sensing baselines, achieving state-of-the-art performance in both point and probabilistic prediction tasks. The model successfully reproduces vegetation dynamics across four future climate scenarios in Europe, with summer 2022 simulations over France aligning closely with known temperature–moisture sensitivities, thereby demonstrating strong spatiotemporal generalization capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

Aug 08, 2026

Existing cross-modal medical image translation methods typically rely on 2D slices or 3D local patches and require separate model training for each task, resulting in limited generalization. This work proposes a whole-volume representation learning framework based on a 3D variational autoencoder, formulating cross-modal translation as a conditional flow-matching problem in latent space. By integrating resolution-aware sampling with multi-task joint training, the approach enables a single model to support diverse modality conversions. It achieves, for the first time, unified whole-volume, multi-task translation with less than 0.15 SSIM performance drop on unseen anatomical regions. The method further supports zero-shot anatomical generalization and unsupervised cross-dataset compositional translation, matching the performance of task-specific models.

0 citationsRead paper

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

Aug 06, 2026

This study investigates whether self-supervised pre-training (SPT) can effectively enhance the diagnostic performance of Transformer models on multimodal, multivariate, and univariate medical time-series data, particularly in data-scarce clinical settings. The work proposes a general-purpose approach that requires no task-specific architectural modifications and incorporates four masking strategies to facilitate representation learning across both modalities and temporal dimensions. Experimental results across three medical time-series tasks demonstrate that SPT improves classification accuracy by 0–6 percentage points, with deeper models exhibiting more pronounced gains. Notably, this is the first study to validate the efficacy of SPT on univariate medical time-series data, demonstrating its strong generalizability, scalability, and robustness under limited-data conditions.

0 citationsRead paper

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Jun 29, 2026

This study addresses the limitations of current radiology report generation models, which rely on holistic evaluation metrics that fail to verify whether diagnoses are grounded in genuine pathological visual evidence, rendering them susceptible to spurious correlations or prior biases. To tackle this issue, the authors introduce the SHOVIR benchmark, incorporating region-level CheXpert labels on MIMIC-CXR and PadChest-GR datasets, and design image- and disease-level occlusion experiments to systematically distinguish between “direct shortcuts” and “contextual shortcuts”—two failure modes of visual dependency. Through a spatially aligned evaluation framework, the work exposes a critical blind spot in existing methods: their neglect of fine-grained regional perception. Experiments across eight state-of-the-art vision-language models reveal that fluent report generation does not necessarily indicate reliable visual grounding, as top-performing models may still exhibit shallow reliance on image evidence.

0 citationsRead paper

VegSim: A Geospatial World Model for Scenario-Conditioned Vegetation Simulation

Jun 20, 2026

Existing vegetation prediction models are constrained by fixed observational meteorological trajectories, limiting their capacity for multi-scenario response analysis. This work proposes the first geospatial world model, which integrates sparse NDVI time series, historical meteorological covariates, and static geographic context through a recurrent latent dynamics architecture. Without requiring scenario-specific supervision, the model simultaneously enables probabilistic forecasting under observed conditions and conditional simulation under user-defined meteorological forcings. Evaluated on the GreenEarthNet benchmark, it significantly outperforms both temporal and remote sensing baselines, achieving state-of-the-art performance in both point and probabilistic prediction tasks. The model successfully reproduces vegetation dynamics across four future climate scenarios in Europe, with summer 2022 simulations over France aligning closely with known temperature–moisture sensitivities, thereby demonstrating strong spatiotemporal generalization capabilities.

0 citationsRead paper

Cross Modality Image Translation In Medical Imaging Using Generative Frameworks

May 13, 2026

This study addresses the lack of standardized evaluation, insufficient clinical validation, and predominant reliance on 2D slices in cross-modal medical image synthesis by introducing the first unified 3D cross-modality image translation assessment framework. The framework incorporates standardized preprocessing, multi-center data partitioning, 3D inference, and multi-level quantitative and clinical evaluations. Seven generative models—including Pix2Pix, CycleGAN, SRGAN, and four implicit models—were systematically benchmarked across multiple oncological imaging tasks. Results indicate that SRGAN performs best overall, yet all methods exhibit limited capability in synthesizing small lesions. In CT-to-PET translation, lesion morphology is better preserved than absolute uptake values. A clinician-involved visual Turing test achieved only 56.7% accuracy in distinguishing real from synthetic images, demonstrating the high clinical realism of the generated outputs.

0 citationsRead paper