Institution profile

Goertek Robotics Co., Ltd.

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

FluSplat: Sparse-View 3D Editing without Test-Time Optimization

Apr 21, 2026

Existing sparse-view 3D editing methods rely on test-time iterative optimization, resulting in high computational costs, cross-view inconsistency, and limited generalization. This work proposes a feed-forward 3D editing framework that eliminates the need for per-scene optimization at test time by incorporating cross-view image-domain regularization and geometric alignment constraints during training. Leveraging text-guided editing, multi-view joint supervision, and a 3D Gaussian splatting representation, the method generates consistent and high-fidelity 3D content without scene-specific refinement. The approach significantly improves cross-view consistency, achieves inference speeds several orders of magnitude faster than existing methods, and maintains high editing fidelity.

0 citationsRead paper

FAST3DIS: Feed-forward Anchored Scene Transformer for 3D Instance Segmentation

Mar 26, 2026

This work proposes an end-to-end feedforward anchored scene Transformer architecture to address the limitations of existing 3D instance segmentation methods, which predominantly rely on non-end-to-end “lift-and-cluster” paradigms that decouple representation learning from segmentation objectives and hinder scalability. The proposed approach introduces learnable 3D anchor generation and anchor-sampling cross-attention mechanisms to achieve multi-view consistent instance segmentation without post-hoc clustering. To mitigate query conflicts and enhance boundary precision, it incorporates dual-level regularization, multi-view contrastive learning, and a dynamic spatial overlap penalty. Evaluated on complex indoor scene datasets, the method significantly outperforms current clustering-based baselines in segmentation accuracy, memory efficiency, and inference speed.

0 citationsRead paper

GRLoc: Geometric Representation Regression for Visual Localization

Nov 17, 2025

In visual localization, Absolute Pose Regression (APR) models often suffer from limited generalization due to end-to-end black-box learning that lacks explicit 3D geometric understanding. To address this, we propose Geometric Representation Regression (GRR), a novel paradigm that abandons direct regression of 6-DoF poses. Instead, GRR separately regresses ray direction bundles (encoding rotation) and point maps (encoding translation), and integrates a differentiable deterministic geometric solver for end-to-end joint optimization in the world coordinate system. Crucially, GRR is the first method to leverage the inverse process of novel view synthesis for pose estimation—explicitly decoupling geometric representation learning from pose solving while embedding strong geometric priors. Evaluated on 7-Scenes and Cambridge Landmarks, GRR achieves state-of-the-art performance, delivering significant improvements in both absolute pose accuracy and cross-scene robustness.

0 citationsRead paper
Recent publications

Latest Papers

FluSplat: Sparse-View 3D Editing without Test-Time Optimization

Apr 21, 2026

Existing sparse-view 3D editing methods rely on test-time iterative optimization, resulting in high computational costs, cross-view inconsistency, and limited generalization. This work proposes a feed-forward 3D editing framework that eliminates the need for per-scene optimization at test time by incorporating cross-view image-domain regularization and geometric alignment constraints during training. Leveraging text-guided editing, multi-view joint supervision, and a 3D Gaussian splatting representation, the method generates consistent and high-fidelity 3D content without scene-specific refinement. The approach significantly improves cross-view consistency, achieves inference speeds several orders of magnitude faster than existing methods, and maintains high editing fidelity.

0 citationsRead paper

FAST3DIS: Feed-forward Anchored Scene Transformer for 3D Instance Segmentation

Mar 26, 2026

This work proposes an end-to-end feedforward anchored scene Transformer architecture to address the limitations of existing 3D instance segmentation methods, which predominantly rely on non-end-to-end “lift-and-cluster” paradigms that decouple representation learning from segmentation objectives and hinder scalability. The proposed approach introduces learnable 3D anchor generation and anchor-sampling cross-attention mechanisms to achieve multi-view consistent instance segmentation without post-hoc clustering. To mitigate query conflicts and enhance boundary precision, it incorporates dual-level regularization, multi-view contrastive learning, and a dynamic spatial overlap penalty. Evaluated on complex indoor scene datasets, the method significantly outperforms current clustering-based baselines in segmentation accuracy, memory efficiency, and inference speed.

0 citationsRead paper

GRLoc: Geometric Representation Regression for Visual Localization

Nov 17, 2025

In visual localization, Absolute Pose Regression (APR) models often suffer from limited generalization due to end-to-end black-box learning that lacks explicit 3D geometric understanding. To address this, we propose Geometric Representation Regression (GRR), a novel paradigm that abandons direct regression of 6-DoF poses. Instead, GRR separately regresses ray direction bundles (encoding rotation) and point maps (encoding translation), and integrates a differentiable deterministic geometric solver for end-to-end joint optimization in the world coordinate system. Crucially, GRR is the first method to leverage the inverse process of novel view synthesis for pose estimation—explicitly decoupling geometric representation learning from pose solving while embedding strong geometric priors. Evaluated on 7-Scenes and Cambridge Landmarks, GRR achieves state-of-the-art performance, delivering significant improvements in both absolute pose accuracy and cross-scene robustness.

0 citationsRead paper