Institution profile

Université Paris-Saclay

Academic institutioneurope · fr
Official website
Research library920linked papers
Opportunities0open roles
Selected work

Representative Papers

A Systematic Comparison of Syntactic Representations of Dependency Parsing

May 29, 2017UDW@NoDaLiDa

This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.

6 citationsRead paper

Variants of Higher-Dimensional Automata

Jan 24, 2026

Higher-dimensional automata (HDA) are often too rigid in their formalism, limiting their direct applicability in modeling and reasoning. This work systematically integrates several weakened variants—such as HDA with interfaces, partial HDA, ST-automata, and relational HDA—and demonstrates, through formal language theory, automata transformations, and algebraic analysis, that these variants fall into only two distinct classes at the language level: those closed under inclusion and those that are not. The paper’s core contributions include the first proof that partial HDA satisfy a Kleene theorem and admit determinization, alongside the establishment of a unified framework that clarifies the expressive power boundaries among all considered variants. These results lay the foundational groundwork for regular expression characterizations and determinization procedures for partial HDA.

1 citationsRead paper

State of the Art of LLM-Enabled Interaction with Visualization

Jan 21, 2026

This study addresses key challenges in the integration of large language models (LLMs) with visualizations—namely, difficulties in multimodal fusion, limited spatial reasoning capabilities, and the absence of standardized evaluation frameworks. Guided by PRISMA guidelines, the authors conduct a systematic literature review of 48 relevant studies and propose a novel six-dimensional taxonomy encompassing application domains, visualization tasks, interaction modalities, and more. This work presents the first comprehensive classification system specifically tailored to LLM–visualization interaction, elucidating prevalent integration patterns of LLMs in data querying, generation, explanation, and navigation. It further synthesizes dominant design paradigms and identifies critical research gaps, particularly concerning accessibility and contextual understanding, thereby establishing a theoretical foundation and future directions for evaluating and developing intelligent, conversational visualization systems.

1 citationsRead paper

Bipartite Tur'an number of paths and other trees

Nov 10, 2025

This paper investigates the maximum number of edges in a bipartite connected graph containing a longest path of prescribed length as a subgraph, with fixed partite set sizes |A| and |B|. Prior work established exact results only for the symmetric case (|A| = |B|) and path lengths at most five. We provide, for the first time, an exact closed-form expression for the extremal number across all partite sizes, establishing a tight functional relationship between the edge bound, path length, and partite set cardinalities. This fully resolves the bipartite path Turán number problem posed by Caro–Patkós–Tuza. Our approach integrates extremal graph theory, structural characterization, and inductive reasoning; we construct extremal graphs and prove their optimality. Additionally, we explore generalizations to star-like trees and other specific tree families, thereby introducing a new paradigm for bipartite extremal problems under subgraph constraints.

1 citationsRead paper

Revisiting the attacker's knowledge in inference attacks against Searchable Symmetric Encryption

Apr 14, 2025

This work investigates the dependence of inference attacks in Searchable Symmetric Encryption (SSE) on the quality of “similar data” available to the adversary. We propose the first general statistical analysis framework that formally defines “similar data” and reveals how its non-uniqueness critically impacts attack robustness. We prove that index size constraints significantly degrade inference attack efficacy and derive a provably secure lower bound on the required index size. Within the leakage-abuse model, we integrate probabilistic modeling with statistical estimation theory and empirically validate our findings on the Enron dataset: imposing an index size cap of 200 reduces the optimal inference attack’s accuracy to below 5% with high probability. Our results yield the first quantifiable, data-similarity-aware defense configuration guideline for SSE systems—bridging theoretical security guarantees with practical deployment constraints.

1 citationsRead paper
Recent publications

Latest Papers

Age of Incorrect Information for Pull-Based State Estimation of General Markov Sources

Aug 13, 2026

This work addresses the challenge of balancing information freshness and correctness in pull-based remote state estimation by formulating a Markov decision process with discounted cost, using the Age of Incorrect Information (AoII) as the optimization objective. By revealing that the belief state depends only on the most recent successful observation and the subsequent duration without updates, the authors propose belief compression, truncation-based approximation, and a hybrid estimator, along with an early steady-switching mechanism. The framework is extended to multi-source settings, where indexability conditions are established and a Whittle index–based heuristic policy is derived. Theoretical and empirical results demonstrate that the optimal single-source policy exhibits a lookup-table waiting structure, while the proposed multi-source scheduling heuristic achieves near-optimal performance with significantly reduced computational overhead and provides computable performance bounds.

0 citationsRead paper

More Than 63% of IEEE VIS Research Liable to be Retracted?! Ethics Approval Statements Protect Participants (and Researchers!)

Aug 13, 2026

This study addresses the widespread omission of ethical approval and informed consent disclosures in IEEE VIS research, which undermines research reproducibility and jeopardizes the rights of both participants and researchers. Conducting the first large-scale systematic content analysis of 255 TVCG-published VIS papers, this work employs manual coding and quantitative analysis to reveal that only 3.2% of studies involving human participants fully reported both ethical approval and informed consent. The findings expose a critical lack of ethical transparency in the VIS community, underscoring the urgent need for standardized ethical reporting guidelines. By providing empirical evidence and foundational insights, this research offers a crucial basis for advancing ethical infrastructure within the discipline.

0 citationsRead paper

Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates

Aug 11, 2026

This work addresses the challenge of inverse parameter calibration in computational physics, where existing surrogate models often lack end-to-end differentiability and physical awareness, limiting their effectiveness. To overcome this, the authors propose a physics-informed latent-space framework based on an autoencoder architecture. The approach enables offline training of a differentiable surrogate model under observable supervision, mapping physical parameters to flow field predictions while embedding variational calibration directly in the latent space. By seamlessly integrating physical constraints with data-driven learning, the method achieves fully end-to-end differentiable surrogate modeling—a first in this domain. Evaluations on two computational fluid dynamics benchmarks demonstrate that, under realistic conditions including noise, low resolution, and partial observability, the proposed framework significantly reduces both calibration error and solution variability.

0 citationsRead paper

Information Bottleneck under Perfect Privacy

Aug 11, 2026

This work addresses the information bottleneck problem under stringent privacy constraints, aiming to construct a compressed representation that preserves utility-relevant information while achieving statistical independence from sensitive variables. To this end, statistical independence is explicitly incorporated as a hard constraint into the information bottleneck framework, and a tailored optimization algorithm based on the Alternating Direction Method of Multipliers (ADMM) is proposed. Under mild regularity conditions, the authors establish the global convergence of the generated iterates, characterize the convergence rate, and extend the analysis to the setting of inexact block updates. These results provide both theoretical guarantees and a practical computational approach for representation learning under perfect privacy.

0 citationsRead paper

HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

Aug 11, 2026

This work addresses the challenge of enabling robots to proactively anticipate human intentions for natural human-robot interaction in real-world environments. To this end, the authors introduce HUI360, the first large-scale in-the-wild 360° egocentric human-robot interaction dataset, accompanied by an efficient annotation pipeline that combines automated labeling with manual correction. The dataset comprises over one million high-quality interaction annotations and is further expanded to six million labels in the SSUP-HRI benchmark. Integrating 360° panoramic video capture, 2D human pose estimation, facial landmark detection, and instance segmentation, the study establishes an end-to-end framework for interaction recognition and intention prediction. The authors publicly release the dataset, code, and baseline models, significantly advancing research and evaluation in human intention anticipation for human-robot interaction.

0 citationsRead paper