Institution profile

AIOZ

Industry researchasia · sg
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Efficient Human-Contact Representation for Human-Scene Interaction

Aug 10, 2026

This work addresses the challenges of high-dimensional input redundancy and low computational efficiency in existing methods for modeling human-scene interaction. To overcome these limitations, the authors propose an efficient representation based on sparse contact masks that retain only the critical contact regions, along with dedicated sparse neural network operators designed to replace conventional dense operations. This approach substantially reduces data redundancy while achieving state-of-the-art reconstruction accuracy across three public benchmarks. Moreover, it delivers a computational speedup of at least 12× compared to baseline methods, effectively balancing reconstruction fidelity with computational efficiency.

0 citationsRead paper

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

Mar 29, 2026

Accurately identifying object affordances in 3D scenes is challenging due to the difficulty of effectively integrating object- and scene-level semantic information. To address this, this work introduces AffordBridge, the first large-scale scene-level dataset comprising 685 high-resolution indoor scenes with aligned RGB images and point cloud annotations. The authors further propose AffordMatcher, a novel method that establishes cross-modal instance correspondences via visual identifiers to enable semantic matching and affordance reasoning between images and point clouds. Experimental results demonstrate that AffordMatcher significantly outperforms existing approaches on the proposed dataset, validating its effectiveness and precision in 3D affordance recognition.

0 citationsRead paper

AeroScene: Progressive Scene Synthesis for Aerial Robotics

Mar 24, 2026

This work addresses the limitations of existing drone simulation frameworks, which rely on manually crafted 3D environments that are difficult to scale and often lack physical plausibility and semantic consistency. To overcome these challenges, we propose the first hierarchical diffusion generative model tailored for aerial robotics tasks. By integrating hierarchy-aware tokenization with multi-branch feature extraction, our approach jointly models global scene layout and local geometric details, enabling progressive 3D scene synthesis. Notably, we introduce a hierarchical diffusion mechanism into drone-centric scene generation and couple it with a physics engine to ensure the generated environments are physically valid and directly usable for downstream tasks such as navigation and landing. Experiments on both a newly curated dataset and established benchmarks demonstrate that our method substantially outperforms existing approaches, successfully generating over 1,000 high-fidelity, physics-ready 3D scenes and significantly enhancing drone navigation performance.

0 citationsRead paper

Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving

Jun 30, 2025

Single-frame visual inputs exhibit poor robustness in complex scenes, while state-of-the-art temporal models suffer from excessive computational overhead, hindering their deployment in federated learning (FL) settings. Method: We propose a lightweight Temporal Transformer decomposition framework that factorizes the global attention matrix into low-rank components, drastically reducing parameter count and computational complexity. We further design an FL-aware distributed training strategy enabling efficient parameter aggregation and real-time inference. The model jointly processes multi-frame images and steering sequences to achieve accurate temporal modeling and cross-modal feature fusion under resource constraints. Results: Our approach outperforms existing SOTA methods on three benchmark datasets, achieving significant accuracy gains and inference latency below 30 ms. Extensive real-world robotic experiments validate its practical deployability and strong generalization capability in heterogeneous edge environments.

0 citationsRead paper

FedEFM: Federated Endovascular Foundation Model with Unseen Data

Jan 28, 2025

To address the dual challenges of scarce annotations and non-shareable cross-institutional data in X-ray catheter/guidewire segmentation for endovascular interventions, this paper proposes the first foundation model framework tailored for vascular interventional imaging under a federated learning paradigm. Methodologically, it integrates federated learning with foundation model pretraining and introduces a novel differentiable Earth Mover’s Distance–based knowledge distillation mechanism to mitigate representation degradation caused by client-wise local data distribution shifts. Evaluated on multiple downstream few-shot segmentation tasks, the framework achieves state-of-the-art performance, significantly improving fine-tuning efficiency and generalization across institutions. This work establishes a practical, scalable, and privacy-preserving paradigm for AI modeling in sensitive medical imaging applications.

0 citationsRead paper
Recent publications

Latest Papers

Efficient Human-Contact Representation for Human-Scene Interaction

Aug 10, 2026

This work addresses the challenges of high-dimensional input redundancy and low computational efficiency in existing methods for modeling human-scene interaction. To overcome these limitations, the authors propose an efficient representation based on sparse contact masks that retain only the critical contact regions, along with dedicated sparse neural network operators designed to replace conventional dense operations. This approach substantially reduces data redundancy while achieving state-of-the-art reconstruction accuracy across three public benchmarks. Moreover, it delivers a computational speedup of at least 12× compared to baseline methods, effectively balancing reconstruction fidelity with computational efficiency.

0 citationsRead paper

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

Mar 29, 2026

Accurately identifying object affordances in 3D scenes is challenging due to the difficulty of effectively integrating object- and scene-level semantic information. To address this, this work introduces AffordBridge, the first large-scale scene-level dataset comprising 685 high-resolution indoor scenes with aligned RGB images and point cloud annotations. The authors further propose AffordMatcher, a novel method that establishes cross-modal instance correspondences via visual identifiers to enable semantic matching and affordance reasoning between images and point clouds. Experimental results demonstrate that AffordMatcher significantly outperforms existing approaches on the proposed dataset, validating its effectiveness and precision in 3D affordance recognition.

0 citationsRead paper

AeroScene: Progressive Scene Synthesis for Aerial Robotics

Mar 24, 2026

This work addresses the limitations of existing drone simulation frameworks, which rely on manually crafted 3D environments that are difficult to scale and often lack physical plausibility and semantic consistency. To overcome these challenges, we propose the first hierarchical diffusion generative model tailored for aerial robotics tasks. By integrating hierarchy-aware tokenization with multi-branch feature extraction, our approach jointly models global scene layout and local geometric details, enabling progressive 3D scene synthesis. Notably, we introduce a hierarchical diffusion mechanism into drone-centric scene generation and couple it with a physics engine to ensure the generated environments are physically valid and directly usable for downstream tasks such as navigation and landing. Experiments on both a newly curated dataset and established benchmarks demonstrate that our method substantially outperforms existing approaches, successfully generating over 1,000 high-fidelity, physics-ready 3D scenes and significantly enhancing drone navigation performance.

0 citationsRead paper

Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving

Jun 30, 2025

Single-frame visual inputs exhibit poor robustness in complex scenes, while state-of-the-art temporal models suffer from excessive computational overhead, hindering their deployment in federated learning (FL) settings. Method: We propose a lightweight Temporal Transformer decomposition framework that factorizes the global attention matrix into low-rank components, drastically reducing parameter count and computational complexity. We further design an FL-aware distributed training strategy enabling efficient parameter aggregation and real-time inference. The model jointly processes multi-frame images and steering sequences to achieve accurate temporal modeling and cross-modal feature fusion under resource constraints. Results: Our approach outperforms existing SOTA methods on three benchmark datasets, achieving significant accuracy gains and inference latency below 30 ms. Extensive real-world robotic experiments validate its practical deployability and strong generalization capability in heterogeneous edge environments.

0 citationsRead paper

FedEFM: Federated Endovascular Foundation Model with Unseen Data

Jan 28, 2025

To address the dual challenges of scarce annotations and non-shareable cross-institutional data in X-ray catheter/guidewire segmentation for endovascular interventions, this paper proposes the first foundation model framework tailored for vascular interventional imaging under a federated learning paradigm. Methodologically, it integrates federated learning with foundation model pretraining and introduces a novel differentiable Earth Mover’s Distance–based knowledge distillation mechanism to mitigate representation degradation caused by client-wise local data distribution shifts. Evaluated on multiple downstream few-shot segmentation tasks, the framework achieves state-of-the-art performance, significantly improving fine-tuning efficiency and generalization across institutions. This work establishes a practical, scalable, and privacy-preserving paradigm for AI modeling in sensitive medical imaging applications.

0 citationsRead paper