Institution profile

Scuola Internazionale Superiore di Studi Avanzati

Academic institutioneurope · it
Official website
Research library81linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond the noise: intrinsic dimension estimation with optimal neighbourhood identification

May 24, 2024arXiv.org

Estimating intrinsic dimensionality (ID) from real-world data is highly sensitive to neighborhood scale: small scales overestimate ID due to noise, while large scales introduce bias from manifold curvature and topology. This work proposes a self-consistent scale selection protocol that identifies the optimal “sweet spot” for ID estimation by enforcing local density constancy. Our key contribution is the first formal coupling of ID estimation and scale selection, resolved via iterative optimization that yields a theoretically guaranteed robust decoupling—effectively suppressing both noise and curvature effects. The method integrates local neighborhood graph construction, asymptotic statistical analysis, and rigorous error-bound derivation. Evaluated on diverse synthetic and real-world datasets, it reduces ID estimation error by over 30% compared to state-of-the-art methods, while significantly improving stability and noise robustness.

3 citationsRead paper

Differential syntactic and semantic encoding in LLMs

Jan 08, 2026arXiv.org

This study investigates the encoding mechanisms and distinctions between syntactic and semantic information in the internal representations of large language models. By computing centroids—mean hidden representations—of sentences sharing either syntactic structure or semantic content, and combining vector subtraction with cross-layer similarity analyses, the work reveals for the first time a partially disentangled linear encoding pattern for syntax and semantics within the model. The findings demonstrate that syntactic and semantic centroids significantly influence the vector similarity of corresponding sentences, and that their encoding trajectories diverge markedly across model layers. This layer-wise differentiation suggests that syntactic and semantic information is organized in distinct ways within deep contextual representations.

2 citationsRead paper

Deriving Neural Scaling Laws from the statistics of natural language

Feb 07, 2026

Existing theoretical frameworks struggle to quantitatively predict neural scaling law exponents from the statistical properties of natural language, particularly in data-constrained regimes. This work proposes the first ab initio theory that requires no free parameters, grounded in information theory and statistical language modeling. By analyzing the decay of token-pair correlations with temporal separation and the decay of next-token conditional entropy with context length, the theory derives a closed-form expression for scaling laws. Relying solely on intrinsic statistical characteristics of natural language, it accurately predicts scaling exponents under limited data conditions. Experimental validation demonstrates close agreement between theoretical predictions and empirical measurements from scratch-trained GPT-2 and LLaMA models on the TinyStories and WikiText datasets.

1 citationsRead paper

Cosmo-FOLD: Fast generation and upscaling of field-level cosmological maps with overlap latent diffusion

Jan 20, 2026

This work proposes Cosmo-FOLD, a novel method that addresses the high computational cost of traditional hydrodynamical cosmological simulations by enabling efficient generation of large-scale, high-fidelity 3D cosmic fields. Leveraging an overlapping latent diffusion model, Cosmo-FOLD accurately upsamples to full-resolution fields using only approximately 1% of the training volume and generalizes across different simulation datasets without fine-tuning. The approach integrates probabilistic diffusion, latent-space modeling, positional encoding, and an overlapping patch strategy to enable efficient 3D field synthesis on a single GPU. Evaluated on the TNG300-2 dataset, the reconstructed dark matter density and gas temperature fields exhibit power spectrum errors below 10% for wavenumbers up to \(k \leq 5 \, h\,\text{Mpc}^{-1}\), while preserving higher-order statistics such as the bispectrum with high fidelity.

1 citationsRead paper
Recent publications

Latest Papers