Institution profile

Universitat Pompeu Fabra

Academic institutioneurope · es
Official website
Research library317linked papers
Opportunities0open roles
Selected work

Representative Papers

NUDF: Neural Unsigned Distance Fields for High Resolution 3D Medical Image Segmentation

Mar 28, 2022IEEE International Symposium on Biomedical Imaging

High-resolution 3D medical image segmentation faces dual challenges of memory bottlenecks and fine-detail loss, especially for topologically complex and morphologically variable structures such as the left atrial appendage. To address this, we propose Neural Unsigned Distance Fields (NUDF), the first method to introduce neural implicit distance fields into medical image segmentation. NUDF employs a coordinate-encoded MLP to directly learn a continuous unsigned distance field from raw CT volumes, thereby avoiding downsampling artifacts and memory constraints inherent to discrete voxel grids. It enables high-fidelity 3D mesh reconstruction with arbitrary topology—including open surfaces—and incorporates continuous distance-based supervision alongside end-to-end differentiable mesh extraction. Evaluated on left atrial appendage segmentation in CT, NUDF achieves sub-voxel accuracy (mean surface error ≈ voxel spacing), significantly outperforming conventional discrete voxel-based methods while reducing memory consumption by an order of magnitude.

4 citations1 influentialRead paper

Linearly Controlled Language Generation with Performative Guarantees

May 24, 2024arXiv.org

To address the need for controllable text generation by large language models (LLMs) in safety-critical applications, this paper tackles the challenge of efficiently ensuring semantic safety of generated outputs. Methodologically, it formalizes semantic constraints as linear structures in the model’s latent space and models the generation process as a trajectory evolution; it then introduces the first gradient-free, closed-form geometric intervention strategy for latent-space control. Theoretically, it establishes the first probabilistic guarantee that generated texts provably reside within a pre-specified safe semantic region. Empirical evaluation on toxicity mitigation demonstrates substantial reduction in harmful content generation while preserving linguistic fluency and lexical diversity—achieving a balanced optimization between controllability and generation quality.

4 citationsRead paper

Beyond the noise: intrinsic dimension estimation with optimal neighbourhood identification

May 24, 2024arXiv.org

Estimating intrinsic dimensionality (ID) from real-world data is highly sensitive to neighborhood scale: small scales overestimate ID due to noise, while large scales introduce bias from manifold curvature and topology. This work proposes a self-consistent scale selection protocol that identifies the optimal “sweet spot” for ID estimation by enforcing local density constancy. Our key contribution is the first formal coupling of ID estimation and scale selection, resolved via iterative optimization that yields a theoretically guaranteed robust decoupling—effectively suppressing both noise and curvature effects. The method integrates local neighborhood graph construction, asymptotic statistical analysis, and rigorous error-bound derivation. Evaluated on diverse synthetic and real-world datasets, it reduces ID estimation error by over 30% compared to state-of-the-art methods, while significantly improving stability and noise robustness.

3 citationsRead paper

The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models

Jan 09, 2026arXiv.org

This study systematically investigates whether and how Transformer-based language models acquire syntactic knowledge. Through a large-scale, systematic literature review synthesizing findings from 337 studies and over 3,000 data points, the work presents the first integrated quantitative assessment of syntactic capabilities across multiple languages and model architectures by combining behavioral experiments, representation probing, and mechanistic interpretability methods. The analysis reveals that Transformers possess substantial syntactic knowledge, yet exhibit limitations in phenomena at the syntax–semantics interface and in low-resource languages. It also highlights a pronounced research bias toward English and BERT-family models, with insufficient coverage of linguistic and architectural diversity. This work provides comprehensive empirical evidence and new directions for understanding the mechanisms and boundaries of syntactic generalization in neural language models.

2 citationsRead paper

Differential syntactic and semantic encoding in LLMs

Jan 08, 2026arXiv.org

This study investigates the encoding mechanisms and distinctions between syntactic and semantic information in the internal representations of large language models. By computing centroids—mean hidden representations—of sentences sharing either syntactic structure or semantic content, and combining vector subtraction with cross-layer similarity analyses, the work reveals for the first time a partially disentangled linear encoding pattern for syntax and semantics within the model. The findings demonstrate that syntactic and semantic centroids significantly influence the vector similarity of corresponding sentences, and that their encoding trajectories diverge markedly across model layers. This layer-wise differentiation suggests that syntactic and semantic information is organized in distinct ways within deep contextual representations.

2 citationsRead paper
Recent publications

Latest Papers