Institution profile

PhiGent Robotics

Industry researchasia · cn
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation

Nov 10, 2025

Existing 4D generation methods suffer from limited spatiotemporal consistency and poor modeling of rapid motion, primarily due to the lack of effective joint spatiotemporal representations. To address this, we propose 4DSTR—a generative network based on 4D Gaussian lattices. Our method introduces two key innovations: (1) a spatiotemporal correction mechanism that explicitly optimizes Gaussian ellipsoid scaling and rotation deformations via time-correlation modeling, ensuring temporal coherence; and (2) an adaptive spatial densification and dynamic pruning strategy that responds in real time to geometric changes induced by abrupt motion. By integrating differentiable 4D rendering with learnable Gaussian point insertion and removal, 4DSTR enables end-to-end video-to-4D generation. Evaluated on standard benchmarks, 4DSTR achieves significant improvements in reconstruction accuracy, spatiotemporal continuity, and robustness to fast motion—setting new state-of-the-art performance.

0 citationsRead paper

OmniNWM: Omniscient Driving Navigation World Models

Oct 21, 2025

Current autonomous driving world models suffer from limitations including single-state modality, short video sequences, coarse-grained action control, and absence of explicit reward modeling—hindering joint modeling of state, action, and reward. This paper introduces the first unified world model for autonomous driving: it enables pixel-level trajectory control via a panoramic Plücker ray representation; constructs a regularized, dense, and differentiable reward function through generative 3D occupancy prediction; and performs multimodal joint modeling over RGB, semantic, depth, and 3D occupancy inputs. The framework supports long-horizon autoregressive generation and closed-loop navigation evaluation. Experiments demonstrate state-of-the-art performance in video fidelity, action accuracy, and long-term stability, while significantly improving simulation capabilities for driving compliance and safety.

0 citationsRead paper

Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition

Oct 18, 2025

Existing gait recognition methods rely on predefined spatial regions and fixed temporal scales, limiting their ability to model the dynamic evolution of motion regions and robustly handle covariate shifts (e.g., viewpoint variations, carried objects). To address these limitations, we propose the Region-aware Dynamic Aggregation and Excitation framework (RDA-RDE), the first approach enabling differentiable automatic discovery of motion-relevant regions, adaptive temporal-scale assignment, and region-specific receptive field modulation. Specifically, the Dynamic Temporal Aggregation module (RDA) learns time-varying motion saliency, while the Region-aware Dynamic Excitation module (RDE) selectively enhances discriminative dynamic regions and suppresses interference-prone static ones—both learned end-to-end. Evaluated on multiple benchmark datasets, RDA-RDE achieves state-of-the-art performance, with particularly significant gains under challenging cross-view and occlusion scenarios involving carried objects.

0 citationsRead paper

SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation

Sep 29, 2025

Text-to-SVG generation faces two key challenges: poor generalization and weak instruction following. To address these, we propose a reasoning-driven instruction alignment framework that explicitly models the visual reasoning process, enabling stepwise generation of complete, editable, and structurally coherent SVG primitives. Our method integrates large language models with multimodal understanding, incorporating staged code generation and supervised fine-tuning. Crucially, we leverage multimodal annotated data to expose and supervise the chain-of-thought reasoning, thereby enhancing reasoning consistency and mitigating hallucination. Experiments demonstrate that our approach significantly outperforms existing methods in generation stability, editability, and visual fidelity—while preserving the inherent advantages of vector graphics—thus advancing the practical deployment of automated graphic design systems.

0 citationsRead paper

DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion

Sep 07, 2025

To address challenges in visual–LiDAR odometry—including sensor misalignment, insufficient temporal information exploitation, and long-sequence error accumulation—this paper proposes a sparse spatiotemporal fusion framework for deep V-LiDAR odometry. The method integrates deep learning with geometric modeling to jointly optimize pose estimation across modalities and time. Key contributions include: (1) a sparse LiDAR query fusion mechanism enabling efficient cross-modal alignment; (2) a temporal interaction update module coupled with prediction-based initialization to enhance inter-frame consistency; and (3) temporal segment-wise training with collective average loss, facilitating global optimization over multiple frames and mitigating scale drift. Evaluated on KITTI and Argoverse benchmarks, the approach achieves state-of-the-art accuracy, significantly reducing pose estimation errors while maintaining real-time inference at 82 ms per frame.

0 citationsRead paper
Recent publications

Latest Papers

4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation

Nov 10, 2025

Existing 4D generation methods suffer from limited spatiotemporal consistency and poor modeling of rapid motion, primarily due to the lack of effective joint spatiotemporal representations. To address this, we propose 4DSTR—a generative network based on 4D Gaussian lattices. Our method introduces two key innovations: (1) a spatiotemporal correction mechanism that explicitly optimizes Gaussian ellipsoid scaling and rotation deformations via time-correlation modeling, ensuring temporal coherence; and (2) an adaptive spatial densification and dynamic pruning strategy that responds in real time to geometric changes induced by abrupt motion. By integrating differentiable 4D rendering with learnable Gaussian point insertion and removal, 4DSTR enables end-to-end video-to-4D generation. Evaluated on standard benchmarks, 4DSTR achieves significant improvements in reconstruction accuracy, spatiotemporal continuity, and robustness to fast motion—setting new state-of-the-art performance.

0 citationsRead paper

OmniNWM: Omniscient Driving Navigation World Models

Oct 21, 2025

Current autonomous driving world models suffer from limitations including single-state modality, short video sequences, coarse-grained action control, and absence of explicit reward modeling—hindering joint modeling of state, action, and reward. This paper introduces the first unified world model for autonomous driving: it enables pixel-level trajectory control via a panoramic Plücker ray representation; constructs a regularized, dense, and differentiable reward function through generative 3D occupancy prediction; and performs multimodal joint modeling over RGB, semantic, depth, and 3D occupancy inputs. The framework supports long-horizon autoregressive generation and closed-loop navigation evaluation. Experiments demonstrate state-of-the-art performance in video fidelity, action accuracy, and long-term stability, while significantly improving simulation capabilities for driving compliance and safety.

0 citationsRead paper

Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition

Oct 18, 2025

Existing gait recognition methods rely on predefined spatial regions and fixed temporal scales, limiting their ability to model the dynamic evolution of motion regions and robustly handle covariate shifts (e.g., viewpoint variations, carried objects). To address these limitations, we propose the Region-aware Dynamic Aggregation and Excitation framework (RDA-RDE), the first approach enabling differentiable automatic discovery of motion-relevant regions, adaptive temporal-scale assignment, and region-specific receptive field modulation. Specifically, the Dynamic Temporal Aggregation module (RDA) learns time-varying motion saliency, while the Region-aware Dynamic Excitation module (RDE) selectively enhances discriminative dynamic regions and suppresses interference-prone static ones—both learned end-to-end. Evaluated on multiple benchmark datasets, RDA-RDE achieves state-of-the-art performance, with particularly significant gains under challenging cross-view and occlusion scenarios involving carried objects.

0 citationsRead paper

SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation

Sep 29, 2025

Text-to-SVG generation faces two key challenges: poor generalization and weak instruction following. To address these, we propose a reasoning-driven instruction alignment framework that explicitly models the visual reasoning process, enabling stepwise generation of complete, editable, and structurally coherent SVG primitives. Our method integrates large language models with multimodal understanding, incorporating staged code generation and supervised fine-tuning. Crucially, we leverage multimodal annotated data to expose and supervise the chain-of-thought reasoning, thereby enhancing reasoning consistency and mitigating hallucination. Experiments demonstrate that our approach significantly outperforms existing methods in generation stability, editability, and visual fidelity—while preserving the inherent advantages of vector graphics—thus advancing the practical deployment of automated graphic design systems.

0 citationsRead paper

DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion

Sep 07, 2025

To address challenges in visual–LiDAR odometry—including sensor misalignment, insufficient temporal information exploitation, and long-sequence error accumulation—this paper proposes a sparse spatiotemporal fusion framework for deep V-LiDAR odometry. The method integrates deep learning with geometric modeling to jointly optimize pose estimation across modalities and time. Key contributions include: (1) a sparse LiDAR query fusion mechanism enabling efficient cross-modal alignment; (2) a temporal interaction update module coupled with prediction-based initialization to enhance inter-frame consistency; and (3) temporal segment-wise training with collective average loss, facilitating global optimization over multiple frames and mitigating scale drift. Evaluated on KITTI and Argoverse benchmarks, the approach achieves state-of-the-art accuracy, significantly reducing pose estimation errors while maintaining real-time inference at 82 ms per frame.

0 citationsRead paper