BinauralVAE: Spatial Audio Reconstruction For World Models
为解决视觉中心世界模型在特定环境下的局限,通过引入BinauralVAE技术,利用变分自编码器架构学习双耳信号的鲁棒潜在表示,增强空间音频重建能力。
为解决视觉中心世界模型在特定环境下的局限,通过引入BinauralVAE技术,利用变分自编码器架构学习双耳信号的鲁棒潜在表示,增强空间音频重建能力。
Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natural language. Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries. To address this issue, Contrastive Language-Image Pretraining (CLIP)-based methods align 3D embeddings with pretrained image and text representations, giving rise to 3D Vision-Language Models (3D VLMs) that support zero-shot classification, cross-modal retrieval, and open-vocabulary recognition of 3D shapes. This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs). We present the main definitions of contrastive learning for multimodal embedding alignment and highlight recent advances in language-guided 3D Gaussian splatting, 3D shape generation, and embodied AI for robotics.
AudioWorldSim通过改进Meta的SoundSpaces 2.0平台,生成逼真的双耳音频数据集,以支持基于音频的机器学习研究。
本文研究了在不满足单交叉条件下,如何通过跳跃或集中处理具有最小有效规模技术的代理问题,并找到全局最优解。
To address low nowcasting accuracy for short-term precipitation in radar-sparse, climatically extreme regions, this paper proposes TUPANN—a physics-aligned neural network that enables high-accuracy, interpretable global precipitation forecasting solely from GOES-16 satellite imagery. Methodologically, TUPANN innovatively incorporates optical flow supervision into a variational encoder-decoder to disentangle motion and intensity evolution; integrates time-varying latent-space modeling with a differentiable advection operator to enhance physical consistency and interpretability; and adopts a MaxViT backbone with multi-city joint training. Experiments across four major climate zones demonstrate that TUPANN achieves state-of-the-art or near-state-of-the-art performance, significantly outperforming baselines for heavy rainfall (≥20 mm/h) prediction. It exhibits strong cross-regional generalization and suitability for near-real-time deployment.
为解决视觉中心世界模型在特定环境下的局限,通过引入BinauralVAE技术,利用变分自编码器架构学习双耳信号的鲁棒潜在表示,增强空间音频重建能力。
Vision-Language Models (VLMs) are reshaping computer vision by aligning visual and textual embeddings, allowing models to recognize visual concepts and reason about them using natural language. Traditional 3D deep-learning models, however, are typically trained for specific tasks, such as classification, segmentation, or detection, and do not naturally support cross-modal retrieval from their embedding spaces using text or images as queries. To address this issue, Contrastive Language-Image Pretraining (CLIP)-based methods align 3D embeddings with pretrained image and text representations, giving rise to 3D Vision-Language Models (3D VLMs) that support zero-shot classification, cross-modal retrieval, and open-vocabulary recognition of 3D shapes. This tutorial provides an overview of 3D VLMs, ranging from basic definitions of 3D representations and their encoding into embeddings to cross-modal contrastive alignment, modern multimodal frameworks, and 3D Vision-Large Language Models (3D VLLMs). We present the main definitions of contrastive learning for multimodal embedding alignment and highlight recent advances in language-guided 3D Gaussian splatting, 3D shape generation, and embodied AI for robotics.
AudioWorldSim通过改进Meta的SoundSpaces 2.0平台,生成逼真的双耳音频数据集,以支持基于音频的机器学习研究。
本文研究了在不满足单交叉条件下,如何通过跳跃或集中处理具有最小有效规模技术的代理问题,并找到全局最优解。
To address low nowcasting accuracy for short-term precipitation in radar-sparse, climatically extreme regions, this paper proposes TUPANN—a physics-aligned neural network that enables high-accuracy, interpretable global precipitation forecasting solely from GOES-16 satellite imagery. Methodologically, TUPANN innovatively incorporates optical flow supervision into a variational encoder-decoder to disentangle motion and intensity evolution; integrates time-varying latent-space modeling with a differentiable advection operator to enhance physical consistency and interpretability; and adopts a MaxViT backbone with multi-city joint training. Experiments across four major climate zones demonstrate that TUPANN achieves state-of-the-art or near-state-of-the-art performance, significantly outperforming baselines for heavy rainfall (≥20 mm/h) prediction. It exhibits strong cross-regional generalization and suitability for near-real-time deployment.