Institution profile

CNRS-Univ. Paris-Dauphine

Academic institutioneurope · fr
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Vanishing L2 regularization for the softmax Multi Armed Bandit

May 05, 2026

This work addresses the long-standing challenge of characterizing the convergence behavior of L2-regularized softmax policy gradient methods as the regularization coefficient vanishes. Focusing on the multi-armed bandit setting, the paper establishes the first rigorous convergence analysis framework for softmax policy gradient under vanishing L2 regularization, circumventing the conventional reliance on convexity assumptions. Through a combination of theoretical analysis and numerical experiments, the study formally proves that the algorithm converges even as the regularization parameter approaches zero. Furthermore, extensive evaluations on standard benchmarks demonstrate that this vanishing-regularization regime consistently outperforms both unregularized and fixed-regularization strategies, highlighting its practical efficacy and theoretical significance.

0 citationsRead paper

KANFormer for Predicting Fill Probabilities via Survival Analysis in Limit Order Books

Dec 05, 2025

This paper addresses the problem of accurately predicting the time-to-fill for limit orders. Conventional approaches rely solely on order book snapshots and neglect the dynamic behavior of market participants. To overcome this limitation, we propose a survival analysis framework that jointly models market state and participant behavior sequences. Our method incorporates Kolmogorov–Arnold Networks to enhance nonlinear representation learning, integrates dilated causal convolution with a Transformer encoder to capture high-dimensional temporal dependencies, and employs SHAP for post-hoc interpretability. Evaluated on CAC 40 futures data, our model achieves statistically significant improvements over state-of-the-art baselines in both calibration (Brier Score) and discriminative performance (Concordance Index). Results demonstrate its effectiveness, robustness, and interpretability in high-frequency trading environments.

0 citationsRead paper
Recent publications

Latest Papers

Vanishing L2 regularization for the softmax Multi Armed Bandit

May 05, 2026

This work addresses the long-standing challenge of characterizing the convergence behavior of L2-regularized softmax policy gradient methods as the regularization coefficient vanishes. Focusing on the multi-armed bandit setting, the paper establishes the first rigorous convergence analysis framework for softmax policy gradient under vanishing L2 regularization, circumventing the conventional reliance on convexity assumptions. Through a combination of theoretical analysis and numerical experiments, the study formally proves that the algorithm converges even as the regularization parameter approaches zero. Furthermore, extensive evaluations on standard benchmarks demonstrate that this vanishing-regularization regime consistently outperforms both unregularized and fixed-regularization strategies, highlighting its practical efficacy and theoretical significance.

0 citationsRead paper

KANFormer for Predicting Fill Probabilities via Survival Analysis in Limit Order Books

Dec 05, 2025

This paper addresses the problem of accurately predicting the time-to-fill for limit orders. Conventional approaches rely solely on order book snapshots and neglect the dynamic behavior of market participants. To overcome this limitation, we propose a survival analysis framework that jointly models market state and participant behavior sequences. Our method incorporates Kolmogorov–Arnold Networks to enhance nonlinear representation learning, integrates dilated causal convolution with a Transformer encoder to capture high-dimensional temporal dependencies, and employs SHAP for post-hoc interpretability. Evaluated on CAC 40 futures data, our model achieves statistically significant improvements over state-of-the-art baselines in both calibration (Brier Score) and discriminative performance (Concordance Index). Results demonstrate its effectiveness, robustness, and interpretability in high-frequency trading environments.

0 citationsRead paper