Institution profile

Kyutai

Research institutioneurope · fr
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

Jun 09, 2026

This work addresses the unnatural interaction behaviors—such as excessively long silences and improper turn-taking—that commonly arise in full-duplex spoken dialogue systems due to their reliance on token-level likelihood optimization. To overcome these limitations, the authors propose a reinforcement learning–based post-training alignment method that systematically encompasses four key interactive dimensions: pause handling, turn-taking, backchannel feedback, and user interruption. The approach introduces an audio segment–level reward function and incorporates a large language model to constrain semantic quality. Experimental results on the Moshi and PersonaPlex benchmarks demonstrate that the method significantly enhances both offline and real-time multi-turn dialogue fluency and naturalness while preserving response accuracy.

0 citationsRead paper

Understanding Data Temporality Impact on Large Language Models Pre-training

May 21, 2026

This work investigates how the temporal ordering of training data affects large language models’ ability to acquire time-sensitive knowledge. Training a 6B-parameter model on chronologically ordered Common Crawl snapshots, the study systematically examines whether sequential exposure to temporally structured corpora mitigates knowledge freezing—a common issue in models trained on shuffled data that impairs their capacity to accurately associate facts with their occurrence times. The authors introduce a novel evaluation benchmark comprising over 7,000 temporally grounded questions and provide the first empirical evidence that temporally ordered pretraining significantly enhances model performance on fact freshness and temporal accuracy, without compromising general language understanding capabilities compared to conventional shuffled-data training.

0 citationsRead paper

ARC-Encoder: learning compressed text representations for large language models

Oct 23, 2025

This work addresses the high computational cost in large language model (LLM) inference caused by contextual redundancy. We propose ARC-Encoder, a general-purpose, architecture- and parameter-agnostic context compression method that requires no modification to the target LLM. ARC-Encoder learns continuous, compact textual representations to replace original token embeddings, directly injecting them into the decoder’s input layer. Its core contribution is a unified encoder architecture—designed for broad compatibility across diverse LLM families—combined with continuous representation learning and a systematic training strategy, enabling 4×–8× context compression without fine-tuning the target model. Experiments demonstrate significant reductions in inference latency and GPU memory consumption across both instruction-tuned and base LLMs, with seamless plug-and-play deployment. ARC-Encoder achieves state-of-the-art performance in efficiency and practicality.

0 citationsRead paper
Recent publications

Latest Papers

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

Jun 09, 2026

This work addresses the unnatural interaction behaviors—such as excessively long silences and improper turn-taking—that commonly arise in full-duplex spoken dialogue systems due to their reliance on token-level likelihood optimization. To overcome these limitations, the authors propose a reinforcement learning–based post-training alignment method that systematically encompasses four key interactive dimensions: pause handling, turn-taking, backchannel feedback, and user interruption. The approach introduces an audio segment–level reward function and incorporates a large language model to constrain semantic quality. Experimental results on the Moshi and PersonaPlex benchmarks demonstrate that the method significantly enhances both offline and real-time multi-turn dialogue fluency and naturalness while preserving response accuracy.

0 citationsRead paper

Understanding Data Temporality Impact on Large Language Models Pre-training

May 21, 2026

This work investigates how the temporal ordering of training data affects large language models’ ability to acquire time-sensitive knowledge. Training a 6B-parameter model on chronologically ordered Common Crawl snapshots, the study systematically examines whether sequential exposure to temporally structured corpora mitigates knowledge freezing—a common issue in models trained on shuffled data that impairs their capacity to accurately associate facts with their occurrence times. The authors introduce a novel evaluation benchmark comprising over 7,000 temporally grounded questions and provide the first empirical evidence that temporally ordered pretraining significantly enhances model performance on fact freshness and temporal accuracy, without compromising general language understanding capabilities compared to conventional shuffled-data training.

0 citationsRead paper

ARC-Encoder: learning compressed text representations for large language models

Oct 23, 2025

This work addresses the high computational cost in large language model (LLM) inference caused by contextual redundancy. We propose ARC-Encoder, a general-purpose, architecture- and parameter-agnostic context compression method that requires no modification to the target LLM. ARC-Encoder learns continuous, compact textual representations to replace original token embeddings, directly injecting them into the decoder’s input layer. Its core contribution is a unified encoder architecture—designed for broad compatibility across diverse LLM families—combined with continuous representation learning and a systematic training strategy, enabling 4×–8× context compression without fine-tuning the target model. Experiments demonstrate significant reductions in inference latency and GPU memory consumption across both instruction-tuned and base LLMs, with seamless plug-and-play deployment. ARC-Encoder achieves state-of-the-art performance in efficiency and practicality.

0 citationsRead paper