Institution profile

Ecole des Hautes Etudes en Sciences Sociales

Academic institutioneurope · fr
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Political attitudes differ but share a common low-dimensional structure across social media and survey data

Mar 02, 2026

This study investigates whether political polarization observed on social media accurately reflects the broader public’s attitudes, focusing on the ideological structure in France within a context of issue misalignment. By integrating large-scale X/Twitter data with nationally representative survey responses and employing dimensionality reduction alongside hierarchical modeling, the research examines how structural factors—such as user activity and visibility—influence the expression of political attitudes. The findings reveal that both online and offline political orientations consistently align along two stable dimensions: “left–right” and “global–local.” Higher user activity is associated with more simplified and polarized attitude structures, while highly visible users exhibit attitudes that closely approximate those of the general public, substantially narrowing the gap between social media discourse and survey-based estimates of public opinion.

0 citationsRead paper

LongTail-Swap: benchmarking language models' abilities on rare words

Oct 05, 2025

This work addresses the weak generalization of language models on lexical long-tail (i.e., rare words). We introduce LongTail-Swap, the first zero-shot evaluation benchmark explicitly designed for the long-tail distribution of pretraining corpora. Built upon the BabyLM dataset, it constructs grammatical/ungrammatical sentence pairs centered on extremely low-frequency tokens; model competence in semantics and syntax for such tokens is assessed via zero-shot average log-probability scores over sentence pairs. Unlike conventional benchmarks emphasizing high-frequency vocabulary, LongTail-Swap systematically exposes severe performance bottlenecks on rare words—revealing that architectural differences yield substantially larger performance gaps in the long tail than in the head. Empirical evaluation across 16 BabyLM models confirms the benchmark’s validity and diagnostic utility for probing long-tail generalization.

0 citationsRead paper
Recent publications

Latest Papers

Political attitudes differ but share a common low-dimensional structure across social media and survey data

Mar 02, 2026

This study investigates whether political polarization observed on social media accurately reflects the broader public’s attitudes, focusing on the ideological structure in France within a context of issue misalignment. By integrating large-scale X/Twitter data with nationally representative survey responses and employing dimensionality reduction alongside hierarchical modeling, the research examines how structural factors—such as user activity and visibility—influence the expression of political attitudes. The findings reveal that both online and offline political orientations consistently align along two stable dimensions: “left–right” and “global–local.” Higher user activity is associated with more simplified and polarized attitude structures, while highly visible users exhibit attitudes that closely approximate those of the general public, substantially narrowing the gap between social media discourse and survey-based estimates of public opinion.

0 citationsRead paper

LongTail-Swap: benchmarking language models' abilities on rare words

Oct 05, 2025

This work addresses the weak generalization of language models on lexical long-tail (i.e., rare words). We introduce LongTail-Swap, the first zero-shot evaluation benchmark explicitly designed for the long-tail distribution of pretraining corpora. Built upon the BabyLM dataset, it constructs grammatical/ungrammatical sentence pairs centered on extremely low-frequency tokens; model competence in semantics and syntax for such tokens is assessed via zero-shot average log-probability scores over sentence pairs. Unlike conventional benchmarks emphasizing high-frequency vocabulary, LongTail-Swap systematically exposes severe performance bottlenecks on rare words—revealing that architectural differences yield substantially larger performance gaps in the long tail than in the head. Empirical evaluation across 16 BabyLM models confirms the benchmark’s validity and diagnostic utility for probing long-tail generalization.

0 citationsRead paper