Institution profile

Universitatea Politehnica Timisoara

Academic institutioneurope · ro
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

Jul 29, 2026

This work addresses the limitations of self-supervised monocular depth estimation, which typically relies on photometric reconstruction loss that entangles depth, pose, and appearance assumptions, thereby constraining representation generalization. For the first time, we introduce the Joint Embedding Predictive Architecture (JEPA) to this task, leveraging a DINOv3-pretrained ViT encoder to predict embeddings of masked target regions from contextual cues in the representation space via a structured masking strategy. The model is jointly optimized with photometric loss and incurs no additional computational overhead at inference. Our approach significantly outperforms baseline methods on KITTI and achieves state-of-the-art or near state-of-the-art zero-shot transfer performance on Make3D and Cityscapes, surpassing leading CNN-based approaches and matching advanced Transformer-based solutions.

0 citationsRead paper

CrashSplat: 2D to 3D Vehicle Damage Segmentation in Gaussian Splatting

Sep 28, 2025

This paper addresses the challenge of accurate 3D modeling of vehicle damage from single-view images. We propose a training-free 2D→3D damage segmentation method. First, camera poses are estimated via Structure-from-Motion. Then, a 3D Gaussian point lattice representing the vehicle geometry is projected onto the image plane; leveraging Z-buffering and a depth-opacity normal distribution model, we perform point cloud filtering and refine the damage mask to achieve geometrically consistent mapping from 2D damage regions to the 3D Gaussian point space. Unlike conventional approaches requiring multi-view consistency, our method precisely localizes fine-grained damage—such as scratches and dents—from a single view, significantly improving geometric accuracy and robustness in 3D reconstruction. Extensive experiments validate its effectiveness, and the implementation is publicly available.

0 citationsRead paper

Quantum-Optimized Selective State Space Model for Efficient Time Series Prediction

Aug 29, 2025

Long-term time series forecasting faces challenges in jointly addressing non-stationarity, multi-scale dependency modeling, and computational efficiency. Existing Transformer-based models (e.g., Autoformer, Informer) suffer from quadratic complexity and degraded long-horizon performance; while state space models (e.g., S-Mamba) achieve linear complexity, they exhibit training instability, sensitivity to initialization, and limited robustness for multivariate settings. This paper proposes the Quantum-optimized Selective State Space Model (Q-SSM), which innovatively incorporates a variational quantum circuit (RY-RX ansatz) as a lightweight gating mechanism. The quantum expectation value adaptively modulates memory updates, preserving O(L) recurrence efficiency while significantly improving training stability and long-range dependency modeling. Evaluated on ETT, Traffic, and Exchange Rate benchmarks, Q-SSM consistently outperforms LSTM, TCN, Reformer, Autoformer, Informer, and S-Mamba—delivering superior accuracy and robustness in multivariate long-horizon forecasting.

0 citationsRead paper

Towards Open World Detection: A Survey

Aug 22, 2025

Computer vision has long been constrained by the closed-world assumption, leading to highly fragmented tasks. Method: This paper introduces “Open-World Detection” (OWD) as the first unified paradigm for general-purpose visual detection in open environments, systematically integrating saliency detection, out-of-distribution detection, zero-shot detection, open-world object detection, and vision-language models. It analyzes task evolution, constructs a coherent knowledge framework, and synthesizes key datasets and methodologies. Contribution/Results: The work reveals an emerging trend toward class-agnostic, scene-adaptive detection fusion. By transcending traditional closed-world constraints, the OWD framework establishes a theoretical foundation and technical roadmap for general visual perception, enabling detection models to robustly generalize beyond narrow, predefined task boundaries into dynamic, open-ended environments.

0 citationsRead paper
Recent publications

Latest Papers

JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation

Jul 29, 2026

This work addresses the limitations of self-supervised monocular depth estimation, which typically relies on photometric reconstruction loss that entangles depth, pose, and appearance assumptions, thereby constraining representation generalization. For the first time, we introduce the Joint Embedding Predictive Architecture (JEPA) to this task, leveraging a DINOv3-pretrained ViT encoder to predict embeddings of masked target regions from contextual cues in the representation space via a structured masking strategy. The model is jointly optimized with photometric loss and incurs no additional computational overhead at inference. Our approach significantly outperforms baseline methods on KITTI and achieves state-of-the-art or near state-of-the-art zero-shot transfer performance on Make3D and Cityscapes, surpassing leading CNN-based approaches and matching advanced Transformer-based solutions.

0 citationsRead paper

CrashSplat: 2D to 3D Vehicle Damage Segmentation in Gaussian Splatting

Sep 28, 2025

This paper addresses the challenge of accurate 3D modeling of vehicle damage from single-view images. We propose a training-free 2D→3D damage segmentation method. First, camera poses are estimated via Structure-from-Motion. Then, a 3D Gaussian point lattice representing the vehicle geometry is projected onto the image plane; leveraging Z-buffering and a depth-opacity normal distribution model, we perform point cloud filtering and refine the damage mask to achieve geometrically consistent mapping from 2D damage regions to the 3D Gaussian point space. Unlike conventional approaches requiring multi-view consistency, our method precisely localizes fine-grained damage—such as scratches and dents—from a single view, significantly improving geometric accuracy and robustness in 3D reconstruction. Extensive experiments validate its effectiveness, and the implementation is publicly available.

0 citationsRead paper

Quantum-Optimized Selective State Space Model for Efficient Time Series Prediction

Aug 29, 2025

Long-term time series forecasting faces challenges in jointly addressing non-stationarity, multi-scale dependency modeling, and computational efficiency. Existing Transformer-based models (e.g., Autoformer, Informer) suffer from quadratic complexity and degraded long-horizon performance; while state space models (e.g., S-Mamba) achieve linear complexity, they exhibit training instability, sensitivity to initialization, and limited robustness for multivariate settings. This paper proposes the Quantum-optimized Selective State Space Model (Q-SSM), which innovatively incorporates a variational quantum circuit (RY-RX ansatz) as a lightweight gating mechanism. The quantum expectation value adaptively modulates memory updates, preserving O(L) recurrence efficiency while significantly improving training stability and long-range dependency modeling. Evaluated on ETT, Traffic, and Exchange Rate benchmarks, Q-SSM consistently outperforms LSTM, TCN, Reformer, Autoformer, Informer, and S-Mamba—delivering superior accuracy and robustness in multivariate long-horizon forecasting.

0 citationsRead paper

Towards Open World Detection: A Survey

Aug 22, 2025

Computer vision has long been constrained by the closed-world assumption, leading to highly fragmented tasks. Method: This paper introduces “Open-World Detection” (OWD) as the first unified paradigm for general-purpose visual detection in open environments, systematically integrating saliency detection, out-of-distribution detection, zero-shot detection, open-world object detection, and vision-language models. It analyzes task evolution, constructs a coherent knowledge framework, and synthesizes key datasets and methodologies. Contribution/Results: The work reveals an emerging trend toward class-agnostic, scene-adaptive detection fusion. By transcending traditional closed-world constraints, the OWD framework establishes a theoretical foundation and technical roadmap for general visual perception, enabling detection models to robustly generalize beyond narrow, predefined task boundaries into dynamic, open-ended environments.

0 citationsRead paper