Institution profile

Carl Zeiss AG

Industry researcheurope · de
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Sterilizable Scene Graph Generation for Operating Rooms

Aug 17, 2026

This study addresses the challenges of parameter redundancy, deployment difficulties, and privacy risks in operating room scene graph generation by proposing the SG-NCA framework. This approach introduces a novel paradigm integrating Neural Cellular Automata (NCA) for scene graph generation and structured representation learning, combined with multi-class segmentation and a lightweight relation predictor. Experimental results demonstrate that SG-NCA achieves performance comparable to mainstream baselines while reducing model parameters by 55 times. Consequently, it enables successful deployment on fanless edge devices, effectively satisfying sterile environment requirements and ensuring data privacy. These findings establish SG-NCA as a viable solution for lightweight medical AI applications, offering a new pathway for secure and efficient intraoperative analysis without compromising accuracy or safety standards in clinical settings.

0 citationsRead paper

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

Jul 17, 2026

This work addresses the challenges in industrial 6D object pose estimation—namely, its strong reliance on texture, requirement for extensive annotated data, and sensitivity to assembly defects—by introducing the first zero-shot framework capable of estimating the pose of unseen objects from RGB images using only textureless 3D models. The method leverages synthetically rendered depth and surface normal maps to align cross-modal features via pretraining, then combines 2D–3D keypoint back-projection, a correspondence filtering mechanism, and a PnP solver to robustly recover pose without object-specific training. Experiments demonstrate state-of-the-art performance on texture-deficient objects, and the authors release a new dataset featuring assembly defects, texture variations, and occlusions to validate real-world applicability.

0 citationsRead paper

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

Jun 22, 2026

Existing video foundation models struggle to effectively model long-horizon procedural videos due to the high computational complexity of self-attention mechanisms and limited ability to distinguish visually similar yet semantically distinct actions. To address these limitations, this work proposes a backbone-agnostic, dense frame-aligned action modeling approach that enables efficient long-video understanding by predicting masked pooled latent vectors. Built upon a predictive joint embedding architecture (P-JEPA) and leveraging features from VJEPA2.1, TSM, and I3D, the model is trained on datasets such as EgoExo4D. It achieves state-of-the-art performance on fine-grained action classification tasks while using an order of magnitude fewer parameters than large language model–based approaches. Furthermore, the method supports real-time streaming inference and temporally precise segmentation.

0 citationsRead paper

OR-Action: Multi-Role Video Understanding with Fine-Grained Actions

Jun 11, 2026

This work addresses the challenges of fine-grained, multi-agent action understanding in operating room videos, where occlusions, clutter, and limited viewpoints hinder effective modeling of long-range temporal structures necessary for coherent action segmentation. To this end, the authors establish the first action-centric benchmark for surgical video understanding, introducing a fine-grained multi-agent action taxonomy and leveraging scene graph state transitions to distill dense action annotations. They propose a purely visual temporal model combined with a multi-view to single-view feature alignment strategy, which significantly enhances action recognition performance in monocular settings without relying on explicit graph structures. Experimental results demonstrate that the proposed approach outperforms existing graph-based methods under both multi-view and single-view evaluation protocols.

0 citationsRead paper

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

May 15, 2026

Current surgical simulation methods struggle to simultaneously achieve visual realism, physically plausible interactions, and out-of-distribution generalization, limiting their clinical applicability. This work proposes SWoMo—a neuro-symbolic world model for cataract surgery—that decouples motion generation from visual rendering: a symbolic component models tool-tissue interaction dynamics using a rule-based simulator and scene graphs, while a diffusion model synthesizes high-fidelity visual appearances. By innovatively integrating symbolic reasoning with neural generation and introducing an inverse pairing strategy to reconstruct real surgical videos into simulation data, the approach enables high-quality sim-to-real transfer and generalization to unseen interaction geometries. Experiments demonstrate that SWoMo significantly outperforms existing methods in visual fidelity, downstream phase detection accuracy, and unsupervised style transfer, validating its clinical relevance and robust generalization capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Sterilizable Scene Graph Generation for Operating Rooms

Aug 17, 2026

This study addresses the challenges of parameter redundancy, deployment difficulties, and privacy risks in operating room scene graph generation by proposing the SG-NCA framework. This approach introduces a novel paradigm integrating Neural Cellular Automata (NCA) for scene graph generation and structured representation learning, combined with multi-class segmentation and a lightweight relation predictor. Experimental results demonstrate that SG-NCA achieves performance comparable to mainstream baselines while reducing model parameters by 55 times. Consequently, it enables successful deployment on fanless edge devices, effectively satisfying sterile environment requirements and ensuring data privacy. These findings establish SG-NCA as a viable solution for lightweight medical AI applications, offering a new pathway for secure and efficient intraoperative analysis without compromising accuracy or safety standards in clinical settings.

0 citationsRead paper

PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

Jul 17, 2026

This work addresses the challenges in industrial 6D object pose estimation—namely, its strong reliance on texture, requirement for extensive annotated data, and sensitivity to assembly defects—by introducing the first zero-shot framework capable of estimating the pose of unseen objects from RGB images using only textureless 3D models. The method leverages synthetically rendered depth and surface normal maps to align cross-modal features via pretraining, then combines 2D–3D keypoint back-projection, a correspondence filtering mechanism, and a PnP solver to robustly recover pose without object-specific training. Experiments demonstrate state-of-the-art performance on texture-deficient objects, and the authors release a new dataset featuring assembly defects, texture variations, and occlusions to validate real-world applicability.

0 citationsRead paper

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

Jun 22, 2026

Existing video foundation models struggle to effectively model long-horizon procedural videos due to the high computational complexity of self-attention mechanisms and limited ability to distinguish visually similar yet semantically distinct actions. To address these limitations, this work proposes a backbone-agnostic, dense frame-aligned action modeling approach that enables efficient long-video understanding by predicting masked pooled latent vectors. Built upon a predictive joint embedding architecture (P-JEPA) and leveraging features from VJEPA2.1, TSM, and I3D, the model is trained on datasets such as EgoExo4D. It achieves state-of-the-art performance on fine-grained action classification tasks while using an order of magnitude fewer parameters than large language model–based approaches. Furthermore, the method supports real-time streaming inference and temporally precise segmentation.

0 citationsRead paper

OR-Action: Multi-Role Video Understanding with Fine-Grained Actions

Jun 11, 2026

This work addresses the challenges of fine-grained, multi-agent action understanding in operating room videos, where occlusions, clutter, and limited viewpoints hinder effective modeling of long-range temporal structures necessary for coherent action segmentation. To this end, the authors establish the first action-centric benchmark for surgical video understanding, introducing a fine-grained multi-agent action taxonomy and leveraging scene graph state transitions to distill dense action annotations. They propose a purely visual temporal model combined with a multi-view to single-view feature alignment strategy, which significantly enhances action recognition performance in monocular settings without relying on explicit graph structures. Experimental results demonstrate that the proposed approach outperforms existing graph-based methods under both multi-view and single-view evaluation protocols.

0 citationsRead paper

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

May 15, 2026

Current surgical simulation methods struggle to simultaneously achieve visual realism, physically plausible interactions, and out-of-distribution generalization, limiting their clinical applicability. This work proposes SWoMo—a neuro-symbolic world model for cataract surgery—that decouples motion generation from visual rendering: a symbolic component models tool-tissue interaction dynamics using a rule-based simulator and scene graphs, while a diffusion model synthesizes high-fidelity visual appearances. By innovatively integrating symbolic reasoning with neural generation and introducing an inverse pairing strategy to reconstruct real surgical videos into simulation data, the approach enables high-quality sim-to-real transfer and generalization to unseen interaction geometries. Experiments demonstrate that SWoMo significantly outperforms existing methods in visual fidelity, downstream phase detection accuracy, and unsupervised style transfer, validating its clinical relevance and robust generalization capabilities.

0 citationsRead paper