An approach with Visual and Tabular Mamba to multimodal medical data using Mixed Fusion
This study addresses the challenge of effectively fusing medical images with clinical and demographic tabular data by proposing a hybrid multimodal fusion approach based on the Mamba architecture. The method employs a Vision Mamba to extract lesion image features and a Tabular Mamba to integrate image-derived prediction probabilities with structured clinical data, enabling end-to-end training for cancer classification. Notably, this work is the first to introduce Mamba into multimodal medical analysis and designs a Mixed Fusion structure that supports SHAP-based interpretability. Evaluated on the NDB-UFES dataset, the proposed approach significantly outperforms Transformer-based baselines while maintaining strong interpretability, demonstrating particularly superior performance in sensitivity-oriented metrics such as recall—making it well-suited for high-stakes clinical diagnostic scenarios.