Institution profile

Instituto Nacional de Matemática Pura e Aplicada

Academic institutionsouthamerica · br
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

An overview of 3D Vision-Language Models

Sep 04, 2026

Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natural language. Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries. To address this issue, Contrastive Language-Image Pretraining (CLIP)-based methods align 3D embeddings with pretrained image and text representations, giving rise to 3D Vision-Language Models (3D VLMs) that support zero-shot classification, cross-modal retrieval, and open-vocabulary recognition of 3D shapes. This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs). We present the main definitions of contrastive learning for multimodal embedding alignment and highlight recent advances in language-guided 3D Gaussian splatting, 3D shape generation, and embodied AI for robotics.

0 citationsRead paper

Precipitation nowcasting of satellite data using physically-aligned neural networks

Nov 07, 2025

To address low nowcasting accuracy for short-term precipitation in radar-sparse, climatically extreme regions, this paper proposes TUPANN—a physics-aligned neural network that enables high-accuracy, interpretable global precipitation forecasting solely from GOES-16 satellite imagery. Methodologically, TUPANN innovatively incorporates optical flow supervision into a variational encoder-decoder to disentangle motion and intensity evolution; integrates time-varying latent-space modeling with a differentiable advection operator to enhance physical consistency and interpretability; and adopts a MaxViT backbone with multi-city joint training. Experiments across four major climate zones demonstrate that TUPANN achieves state-of-the-art or near-state-of-the-art performance, significantly outperforming baselines for heavy rainfall (≥20 mm/h) prediction. It exhibits strong cross-regional generalization and suitability for near-real-time deployment.

0 citationsRead paper
Recent publications

Latest Papers

An overview of 3D Vision-Language Models

Sep 04, 2026

Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natural language. Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries. To address this issue, Contrastive Language-Image Pretraining (CLIP)-based methods align 3D embeddings with pretrained image and text representations, giving rise to 3D Vision-Language Models (3D VLMs) that support zero-shot classification, cross-modal retrieval, and open-vocabulary recognition of 3D shapes. This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs). We present the main definitions of contrastive learning for multimodal embedding alignment and highlight recent advances in language-guided 3D Gaussian splatting, 3D shape generation, and embodied AI for robotics.

0 citationsRead paper

Precipitation nowcasting of satellite data using physically-aligned neural networks

Nov 07, 2025

To address low nowcasting accuracy for short-term precipitation in radar-sparse, climatically extreme regions, this paper proposes TUPANN—a physics-aligned neural network that enables high-accuracy, interpretable global precipitation forecasting solely from GOES-16 satellite imagery. Methodologically, TUPANN innovatively incorporates optical flow supervision into a variational encoder-decoder to disentangle motion and intensity evolution; integrates time-varying latent-space modeling with a differentiable advection operator to enhance physical consistency and interpretability; and adopts a MaxViT backbone with multi-city joint training. Experiments across four major climate zones demonstrate that TUPANN achieves state-of-the-art or near-state-of-the-art performance, significantly outperforming baselines for heavy rainfall (≥20 mm/h) prediction. It exhibits strong cross-regional generalization and suitability for near-real-time deployment.

0 citationsRead paper