Institution profile

ELLIS Institute Finland

Academic institutioneurope · fi
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

May 21, 2026

This study investigates the relationship between the performance of embedding models and the structural properties of their embedding spaces, with the aim of predicting downstream task effectiveness. Leveraging the MTEB benchmark, the authors evaluate 25 prominent embedding models across five tasks in both English and multilingual settings. They characterize the local and linear structures of embedding spaces using nearest-neighbor overlap and independent component analysis (ICA). The work reveals, for the first time, a remarkably high correlation (up to 0.97) between the degree of local structure preservation in embedding spaces and model performance on downstream tasks. Furthermore, it demonstrates that different tasks exhibit distinct dependencies on local versus linear structural information. These findings indicate that structural characteristics of embedding spaces can effectively predict model performance across diverse tasks, including retrieval, bilingual text mining, pair classification, and summarization.

0 citationsRead paper

Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs

Feb 17, 2026

This study addresses the challenge posed by the unstructured nature of large-scale historical interview transcripts—comprising 350,000 activity and organization mentions—which has hindered quantitative research on social integration. To overcome this, the authors develop a participation taxonomy that systematically categorizes activities along dimensions of type, sociability, frequency, and physical demand. For the first time, they apply open-source large language models to historical oral archives, employing a multi-round voting strategy to automatically annotate a human-labeled gold-standard dataset. The model’s outputs demonstrate high agreement with expert judgments, effectively transforming unstructured narrative text into structured social indicators. This approach not only enables scalable quantification of historical social participation but also offers a methodological innovation that expands data resources for social integration research.

0 citationsRead paper
Recent publications

Latest Papers

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

May 21, 2026

This study investigates the relationship between the performance of embedding models and the structural properties of their embedding spaces, with the aim of predicting downstream task effectiveness. Leveraging the MTEB benchmark, the authors evaluate 25 prominent embedding models across five tasks in both English and multilingual settings. They characterize the local and linear structures of embedding spaces using nearest-neighbor overlap and independent component analysis (ICA). The work reveals, for the first time, a remarkably high correlation (up to 0.97) between the degree of local structure preservation in embedding spaces and model performance on downstream tasks. Furthermore, it demonstrates that different tasks exhibit distinct dependencies on local versus linear structural information. These findings indicate that structural characteristics of embedding spaces can effectively predict model performance across diverse tasks, including retrieval, bilingual text mining, pair classification, and summarization.

0 citationsRead paper

Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs

Feb 17, 2026

This study addresses the challenge posed by the unstructured nature of large-scale historical interview transcripts—comprising 350,000 activity and organization mentions—which has hindered quantitative research on social integration. To overcome this, the authors develop a participation taxonomy that systematically categorizes activities along dimensions of type, sociability, frequency, and physical demand. For the first time, they apply open-source large language models to historical oral archives, employing a multi-round voting strategy to automatically annotate a human-labeled gold-standard dataset. The model’s outputs demonstrate high agreement with expert judgments, effectively transforming unstructured narrative text into structured social indicators. This approach not only enables scalable quantification of historical social participation but also offers a methodological innovation that expands data resources for social integration research.

0 citationsRead paper