Institution profile

Mistral AI

Industry researcheurope · fr
Official website
Research library24linked papers
Opportunities0open roles
Selected work

Representative Papers

Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning

Feb 26, 2026

This study addresses the limited reasoning capabilities of current vision-language models (VLMs), which stem from reporting bias in training data that systematically omits implicit information—such as spatial relations, temporal dynamics, negation, and counting. For the first time, the authors formally integrate pragmatic theories of reporting bias into vision-language learning, constructing a targeted evaluation benchmark to assess multiple models, including OpenCLIP, LLaVA-1.5, and Molmo. Their findings reveal that merely scaling up data volume or incorporating multilingual corpora fails to rectify these reasoning gaps. In contrast, explicitly designing and integrating annotations that surface such implicit information substantially enhances model performance. This work challenges the prevailing “scale-is-all-you-need” paradigm, underscoring the necessity of deliberately curating training data that supports robust multimodal reasoning.

1 citationsRead paper

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

Aug 17, 2026

This study addresses the challenges of applying value functions in reinforcement learning for large language models and the inefficiency of Group Relative Policy Optimization (GRPO). We propose a Privileged Value Function coupled with a tethered adaptive interpolation mechanism to optimize credit assignment and enable dynamic baseline adjustment. This approach significantly enhances training stability and sample efficiency. Empirical results demonstrate that our method not only outperforms standard value function baselines in reasoning tasks but also matches or surpasses GRPO performance. By effectively overcoming existing bottlenecks, this work establishes an efficient new paradigm for optimizing the reasoning capabilities of large language models via reinforcement learning.

0 citationsRead paper

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

Aug 04, 2026

This study investigates how the rigid formatting and stylistic constraints inherent in traditional chain-of-thought (CoT) prompting may inadvertently hinder the core reasoning capabilities of large language models. Through systematic evaluation on mathematical reasoning benchmarks such as GSM8K, the authors compare zero-shot soft prompting against few-shot CoT across multiple medium-scale specialized and general-purpose models. They reveal a previously unobserved phenomenon: as model capacity increases, standard CoT prompting becomes a performance bottleneck, whereas lightweight soft prompting in a zero-shot setting consistently outperforms few-shot CoT—evidenced by an accuracy improvement from 77% to 84% on the Mathstral model, with similar gains observed across general models. These findings challenge the prevailing assumption of CoT’s universal efficacy and offer a new direction for designing more efficient reasoning prompts.

0 citationsRead paper

From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models

Jul 29, 2026

This work addresses the lack of spatial uncertainty modeling in YOLO-Pose for keypoint localization. The authors propose a lightweight, post-hoc probabilistic extension that introduces an additional probability head to predict input-dependent 2×2 covariance matrices, enabling calibrated bivariate Gaussian or Student-t distributions over original keypoints. This is the first approach to equip YOLO-Pose with keypoint-level predictive distributions. A novel evaluation protocol is introduced, combining distribution calibration diagnostics with Average Keypoint Precision (AKP). Experiments on COCO demonstrate that the method effectively supports reliability-based keypoint ranking, with the Student-t formulation yielding more accurate residual distribution fitting. Furthermore, in an aircraft visual landing task, the calibrated covariance enables uncertainty-aware pose estimation and sensor fusion.

0 citationsRead paper

Stochastic Counterdiabatic Driving via Biorthogonal Liouvillian Eigenmodes

Jul 27, 2026

This work addresses the bias in free energy estimation arising from nonadiabatic hysteresis in stochastically driven systems under finite-time protocols. The authors propose a canonical transformation framework based on the exact spectral decomposition of the Fokker–Planck generator, which—by leveraging the biorthogonal eigenmodes of the Liouville operator—constructs an antiadiabatic correction field that completely eliminates nonadiabatic hysteresis at arbitrary driving speeds. Formally analogous to Berry’s transitionless quantum driving, this approach achieves zero-variance free energy estimates. Numerical validation in double-well and harmonic potential models demonstrates dramatic improvements: total variation distance and Kullback–Leibler divergence are reduced by approximately 12 and 16 orders of magnitude, respectively, while dissipated work approaches machine precision, confirming the theoretical exactness of the method.

0 citationsRead paper
Recent publications

Latest Papers

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

Aug 17, 2026

This study addresses the challenges of applying value functions in reinforcement learning for large language models and the inefficiency of Group Relative Policy Optimization (GRPO). We propose a Privileged Value Function coupled with a tethered adaptive interpolation mechanism to optimize credit assignment and enable dynamic baseline adjustment. This approach significantly enhances training stability and sample efficiency. Empirical results demonstrate that our method not only outperforms standard value function baselines in reasoning tasks but also matches or surpasses GRPO performance. By effectively overcoming existing bottlenecks, this work establishes an efficient new paradigm for optimizing the reasoning capabilities of large language models via reinforcement learning.

0 citationsRead paper

Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve

Aug 04, 2026

This study investigates how the rigid formatting and stylistic constraints inherent in traditional chain-of-thought (CoT) prompting may inadvertently hinder the core reasoning capabilities of large language models. Through systematic evaluation on mathematical reasoning benchmarks such as GSM8K, the authors compare zero-shot soft prompting against few-shot CoT across multiple medium-scale specialized and general-purpose models. They reveal a previously unobserved phenomenon: as model capacity increases, standard CoT prompting becomes a performance bottleneck, whereas lightweight soft prompting in a zero-shot setting consistently outperforms few-shot CoT—evidenced by an accuracy improvement from 77% to 84% on the Mathstral model, with similar gains observed across general models. These findings challenge the prevailing assumption of CoT’s universal efficacy and offer a new direction for designing more efficient reasoning prompts.

0 citationsRead paper

From Keypoints to Predictive Distributions: Post-Hoc Uncertainty for YOLO-Pose Models

Jul 29, 2026

This work addresses the lack of spatial uncertainty modeling in YOLO-Pose for keypoint localization. The authors propose a lightweight, post-hoc probabilistic extension that introduces an additional probability head to predict input-dependent 2×2 covariance matrices, enabling calibrated bivariate Gaussian or Student-t distributions over original keypoints. This is the first approach to equip YOLO-Pose with keypoint-level predictive distributions. A novel evaluation protocol is introduced, combining distribution calibration diagnostics with Average Keypoint Precision (AKP). Experiments on COCO demonstrate that the method effectively supports reliability-based keypoint ranking, with the Student-t formulation yielding more accurate residual distribution fitting. Furthermore, in an aircraft visual landing task, the calibrated covariance enables uncertainty-aware pose estimation and sensor fusion.

0 citationsRead paper

Stochastic Counterdiabatic Driving via Biorthogonal Liouvillian Eigenmodes

Jul 27, 2026

This work addresses the bias in free energy estimation arising from nonadiabatic hysteresis in stochastically driven systems under finite-time protocols. The authors propose a canonical transformation framework based on the exact spectral decomposition of the Fokker–Planck generator, which—by leveraging the biorthogonal eigenmodes of the Liouville operator—constructs an antiadiabatic correction field that completely eliminates nonadiabatic hysteresis at arbitrary driving speeds. Formally analogous to Berry’s transitionless quantum driving, this approach achieves zero-variance free energy estimates. Numerical validation in double-well and harmonic potential models demonstrates dramatic improvements: total variation distance and Kullback–Leibler divergence are reduced by approximately 12 and 16 orders of magnitude, respectively, while dissipated work approaches machine precision, confirming the theoretical exactness of the method.

0 citationsRead paper

A Shortcut to Statistically Steady-State Turbulence with Flow Matching

Jul 14, 2026

High-fidelity gyrokinetic turbulence simulations are computationally expensive due to the need to resolve the full temporal evolution from transients to statistical steady states, and efficient reduced-order models remain lacking. This work addresses this challenge by leveraging the ergodicity hypothesis to bypass explicit time integration and, for the first time, applies flow matching to directly generate steady-state statistical distributions of saturated turbulence in five-dimensional phase space. Conditioned on dimensionless operational parameters, the proposed latent-space generative model, GyroFlow, synthesizes high-quality steady-state snapshots from noise by integrating a pre-trained physics-informed metric with a conditional generation mechanism. A new evaluation metric, FGyD, is introduced to assess fidelity. Experiments demonstrate that GyroFlow outperforms existing autoregressive, reduced-order, and generative approaches in both sample quality and downstream flux prediction accuracy, significantly accelerating simulation and enabling effective hot-starting of the original solver.

0 citationsRead paper