Institution profile

LINAGORA

Industry researcheurope · fr
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

Aug 03, 2026

This study investigates whether reinforcement learning with verifiable rewards (RLVR) genuinely enhances the reasoning capabilities of large language models or merely improves sampling efficiency. To this end, we introduce BODHI-Trees—a novel tree-based representation that extracts semantically equivalent structures from mathematical reasoning trajectories—and propose semantic branching entropy as a new metric to quantify reasoning diversity. Through controlled maze experiments and trajectory analyses, we find that while RLVR strengthens constraint adherence and backtracking abilities, it substantially contracts the semantic reasoning space, leading to a concurrent collapse in both policy entropy and semantic branching entropy. These findings suggest that the efficiency gains conferred by RLVR may come at the cost of reduced reasoning diversity.

0 citationsRead paper

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?

Jul 15, 2025

This work investigates whether large language models (LLMs) implicitly encode structured causal dependencies underlying mathematical reasoning, and how chain-of-thought (CoT) prompting supports this capability. Method: We propose the Causal Chain Graph (CCG)—a directed acyclic graph that models fine-grained causal mediation and path preferences among reasoning steps—and automatically construct it from model-generated reasoning traces. To enable rigorous analysis, we introduce KisMATH, a benchmark comprising 1,671 math problems annotated with step-level causal structure. Contribution/Results: Through multi-model empirical evaluation across 15 open-source LLMs and graph-alignment interventions, we demonstrate that LLMs consistently follow CCG-structured causal paths during reasoning. This provides the first causally grounded, graph-based interpretability evidence for CoT’s internal mechanism, revealing that mathematical reasoning in LLMs is not merely sequential but governed by latent causal structure—thereby advancing our understanding of the fundamental nature of LLM-based mathematical reasoning.

0 citationsRead paper

LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

Apr 03, 2025

Tunisian Arabic speech recognition faces a low-resource bottleneck due to scarce annotated speech data and inadequate modeling of English–French code-switching. To address this, we introduce LinTO—the first high-quality, open-source ASR dataset for Tunisian Arabic—comprising multi-source text, real-world speech recordings (including natural code-switching), and phoneme-level aligned transcriptions. We propose a dialect-specific phonological annotation schema and an explicit code-switching modeling framework. Leveraging systematic speech collection, multilingual alignment, audio augmentation, and a phonology-informed evaluation methodology, LinTO contains tens of thousands of precisely annotated utterances, filling a critical benchmark resource gap. Evaluated on standard test sets, ASR models trained on LinTO achieve a 22% relative reduction in word error rate (WER). LinTO thus establishes the first authoritative ASR benchmark and end-to-end technical solution for Tunisian Arabic.

0 citationsRead paper

The Lucie-7B LLM and the Lucie Training Dataset: Open resources for multilingual language generation

Mar 15, 2025

This work addresses English-centric bias in multilingual large language models, particularly mitigating data scarcity for French and other European languages. We propose an OSI-compliant, French–English balanced pretraining paradigm and construct the first systematically curated multilingual foundational resource integrating French cultural heritage texts—ensuring both copyright compliance and open reproducibility. We release Lucie-7B, a 7-billion-parameter Transformer-based model pretrained on approximately equal proportions (≈33% each) of French and English data. Following instruction fine-tuning augmented with human-annotated data, Lucie-7B achieves competitive performance against state-of-the-art models on multilingual benchmarks. All model weights, training checkpoints, data processing scripts, and comprehensive documentation are fully open-sourced on Hugging Face and GitHub. This end-to-end open stack advances responsible, reproducible, and culturally inclusive multilingual AI research and development.

0 citationsRead paper
Recent publications

Latest Papers

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

Aug 03, 2026

This study investigates whether reinforcement learning with verifiable rewards (RLVR) genuinely enhances the reasoning capabilities of large language models or merely improves sampling efficiency. To this end, we introduce BODHI-Trees—a novel tree-based representation that extracts semantically equivalent structures from mathematical reasoning trajectories—and propose semantic branching entropy as a new metric to quantify reasoning diversity. Through controlled maze experiments and trajectory analyses, we find that while RLVR strengthens constraint adherence and backtracking abilities, it substantially contracts the semantic reasoning space, leading to a concurrent collapse in both policy entropy and semantic branching entropy. These findings suggest that the efficiency gains conferred by RLVR may come at the cost of reduced reasoning diversity.

0 citationsRead paper

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?

Jul 15, 2025

This work investigates whether large language models (LLMs) implicitly encode structured causal dependencies underlying mathematical reasoning, and how chain-of-thought (CoT) prompting supports this capability. Method: We propose the Causal Chain Graph (CCG)—a directed acyclic graph that models fine-grained causal mediation and path preferences among reasoning steps—and automatically construct it from model-generated reasoning traces. To enable rigorous analysis, we introduce KisMATH, a benchmark comprising 1,671 math problems annotated with step-level causal structure. Contribution/Results: Through multi-model empirical evaluation across 15 open-source LLMs and graph-alignment interventions, we demonstrate that LLMs consistently follow CCG-structured causal paths during reasoning. This provides the first causally grounded, graph-based interpretability evidence for CoT’s internal mechanism, revealing that mathematical reasoning in LLMs is not merely sequential but governed by latent causal structure—thereby advancing our understanding of the fundamental nature of LLM-based mathematical reasoning.

0 citationsRead paper

LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

Apr 03, 2025

Tunisian Arabic speech recognition faces a low-resource bottleneck due to scarce annotated speech data and inadequate modeling of English–French code-switching. To address this, we introduce LinTO—the first high-quality, open-source ASR dataset for Tunisian Arabic—comprising multi-source text, real-world speech recordings (including natural code-switching), and phoneme-level aligned transcriptions. We propose a dialect-specific phonological annotation schema and an explicit code-switching modeling framework. Leveraging systematic speech collection, multilingual alignment, audio augmentation, and a phonology-informed evaluation methodology, LinTO contains tens of thousands of precisely annotated utterances, filling a critical benchmark resource gap. Evaluated on standard test sets, ASR models trained on LinTO achieve a 22% relative reduction in word error rate (WER). LinTO thus establishes the first authoritative ASR benchmark and end-to-end technical solution for Tunisian Arabic.

0 citationsRead paper

The Lucie-7B LLM and the Lucie Training Dataset: Open resources for multilingual language generation

Mar 15, 2025

This work addresses English-centric bias in multilingual large language models, particularly mitigating data scarcity for French and other European languages. We propose an OSI-compliant, French–English balanced pretraining paradigm and construct the first systematically curated multilingual foundational resource integrating French cultural heritage texts—ensuring both copyright compliance and open reproducibility. We release Lucie-7B, a 7-billion-parameter Transformer-based model pretrained on approximately equal proportions (≈33% each) of French and English data. Following instruction fine-tuning augmented with human-annotated data, Lucie-7B achieves competitive performance against state-of-the-art models on multilingual benchmarks. All model weights, training checkpoints, data processing scripts, and comprehensive documentation are fully open-sourced on Hugging Face and GitHub. This end-to-end open stack advances responsible, reproducible, and culturally inclusive multilingual AI research and development.

0 citationsRead paper