Institution profile

Ashesi University

Academic institutionafrica · gh
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

VAMAE: Vessel-Aware Masked Autoencoders for OCT Angiography

Apr 07, 2026

This work addresses the challenge of self-supervised representation learning in OCTA images, where sparse vasculature and strong topological constraints hinder effective feature learning. To this end, the authors propose a vessel-aware masked autoencoder framework that integrates vessel saliency with skeleton priors to devise an anatomy-guided, non-uniform masking strategy. By jointly optimizing multi-objective reconstruction tasks, the method simultaneously preserves vascular appearance, structural continuity, and topological fidelity, thereby enabling geometry-aware learning of vessel connectivity and branching patterns. Experiments on the OCTA-500 benchmark demonstrate that the proposed approach significantly outperforms standard masked autoencoders, with particularly notable gains in label-scarce settings.

0 citationsRead paper

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Oct 20, 2025

African languages—representing a significant portion of the world’s linguistic diversity—are severely underrepresented in multimodal AI, particularly in image captioning, due to scarce annotated data and limited model support. Method: This work introduces the first large-scale vision-to-language framework for 20 African languages. It constructs a semantically aligned, high-quality multilingual image-caption dataset; designs a dynamic quality assurance pipeline integrating context-aware translation, model ensembling (SigLIP + NLLB-200), and adaptive token replacement; and develops a unified, 0.5B-parameter vision-to-text architecture optimized for low-resource settings. Contribution/Results: We release the first open-source, African-language–focused image captioning dataset and corresponding pre-trained models. Our framework establishes a new multilingual generation paradigm that balances accuracy and scalability, achieving substantial performance gains on cross-modal tasks for low-resource languages. This advances inclusive, equitable multimodal AI development and sets a foundation for future research in under-resourced language modalities.

0 citationsRead paper
Recent publications

Latest Papers

VAMAE: Vessel-Aware Masked Autoencoders for OCT Angiography

Apr 07, 2026

This work addresses the challenge of self-supervised representation learning in OCTA images, where sparse vasculature and strong topological constraints hinder effective feature learning. To this end, the authors propose a vessel-aware masked autoencoder framework that integrates vessel saliency with skeleton priors to devise an anatomy-guided, non-uniform masking strategy. By jointly optimizing multi-objective reconstruction tasks, the method simultaneously preserves vascular appearance, structural continuity, and topological fidelity, thereby enabling geometry-aware learning of vessel connectivity and branching patterns. Experiments on the OCTA-500 benchmark demonstrate that the proposed approach significantly outperforms standard masked autoencoders, with particularly notable gains in label-scarce settings.

0 citationsRead paper

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

Oct 20, 2025

African languages—representing a significant portion of the world’s linguistic diversity—are severely underrepresented in multimodal AI, particularly in image captioning, due to scarce annotated data and limited model support. Method: This work introduces the first large-scale vision-to-language framework for 20 African languages. It constructs a semantically aligned, high-quality multilingual image-caption dataset; designs a dynamic quality assurance pipeline integrating context-aware translation, model ensembling (SigLIP + NLLB-200), and adaptive token replacement; and develops a unified, 0.5B-parameter vision-to-text architecture optimized for low-resource settings. Contribution/Results: We release the first open-source, African-language–focused image captioning dataset and corresponding pre-trained models. Our framework establishes a new multilingual generation paradigm that balances accuracy and scalability, achieving substantial performance gains on cross-modal tasks for low-resource languages. This advances inclusive, equitable multimodal AI development and sets a foundation for future research in under-resourced language modalities.

0 citationsRead paper