Institution profile

Steklov Mathematical Institute

Academic institutioneurope · ru
Official website
Research library22linked papers
Opportunities0open roles
Selected work

Representative Papers

Against Opacity: Explainable AI and Large Language Models for Effective Digital Advertising

Oct 26, 2023ACM Multimedia

Digital advertising platforms (e.g., Meta Ads) suffer from algorithmic opacity, hindering advertisers’ understanding of audience targeting, pricing mechanisms, and ad relevance—thereby impeding data-driven decision-making. To address this, we propose SODA: the first explainable advertising analytics framework integrating multimodal text-image models with large language models (LLMs). Our method introduces a natural-language–based interactive explanation interface tailored for non-technical marketing professionals, enabling automated competitive ad summarization, attribution analysis, and click-through rate (CTR) prediction. By synergistically combining eXplainable AI (XAI) techniques with natural language generation and understanding, SODA enhances predictive accuracy while delivering actionable, trustworthy AI-assisted insights. Evaluated in real-world deployment scenarios, SODA significantly improves interpretability without compromising performance, empowering marketers to make informed, auditable decisions grounded in transparent model reasoning.

12 citationsRead paper

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

Jul 30, 2026

This study addresses the lack of quantitative morphological research on Dungan based on authentic corpora. It presents the first systematic construction of a finite-state morphological analyzer (built with HFST) covering dialects from Gansu and Shaanxi, formalizing existing linguistic knowledge and evaluating its lexical coverage and ambiguity distribution across three text genres. The findings reveal that inflectional morphology in Dungan is rare—only 9.3% of tokens bear morphological markers—and 78.1% of lemmas are unambiguous, indicating a highly closed grammatical core with openness primarily at the lexical level. Out-of-vocabulary items account for 78%–95% of analysis failures, and integrating the morphological model improves lemma coverage by 5.2 percentage points.

0 citationsRead paper

First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers

Jul 18, 2026

This study investigates the tension between the predictability of unidirectional perturbations and the fragility of multidirectional combinations in local task adaptation of Transformers. Leveraging multi-task LoRA operating points, the authors systematically evaluate eight local adaptation properties across nine models of varying scales, establishing—via a preregistered protocol—the quantitative boundary between first-order predictability and pairwise combinatorial fragility for the first time. Integrating LoRA fine-tuning, task arithmetic, activation steering, gradient analysis, and Lie bracket modeling, they demonstrate that unidirectional perturbations are highly predictable within a first-order regime (effective window up to 10⁻²), yet over one-third of model–task pairs exhibit order sensitivity at even smaller scales. The Lie bracket precisely captures this ordering dependence, with a median ratio of predicted to empirical deviation of merely 1.002.

0 citationsRead paper

Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration

Jul 18, 2026

This work addresses the limitation of existing large model memory evaluations, which predominantly rely on dialogue or summarization and fail to assess scientific agents’ ability to recover evidence from full research papers. The authors propose the first full-text–oriented scientific memory evaluation paradigm, introducing two new benchmarks—PAIM and PTr—that formalize scientific memory as a budget-constrained, modality-aware context recovery task. Their approach combines sparse (BM25) and dense retrieval in a hybrid strategy, incorporating multi-granularity text ingestion, preservation of original content, controlled retrieval budgets, and consistency calibration across multiple evaluators, including humans. Experiments reveal that under strict retrieval budget constraints, the performance advantage of leading systems vanishes; hybrid retrieval substantially improves results; and LLM-based evaluators align closely with human judgments, achieving a resolution of 0.1 points.

0 citationsRead paper

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

Jul 07, 2026

Existing evaluations of code agents predominantly focus on task success or failure, overlooking the nuanced interaction trajectories observed in real-world usage. This work proposes the first fine-grained evaluation framework that integrates formal verification with large language model–generated commentary, leveraging complete interaction traces from actual production environments to establish a multidimensional and interpretable assessment mechanism. The framework enables detailed model behavior diagnostics, supports regression detection across model versions, and incorporates an automated evaluation pipeline. Its effectiveness has been validated through internal agent iterations and nightly regression testing, and the associated benchmark suite has been open-sourced to provide the research community with a high-value evaluation tool.

0 citationsRead paper
Recent publications

Latest Papers

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

Jul 30, 2026

This study addresses the lack of quantitative morphological research on Dungan based on authentic corpora. It presents the first systematic construction of a finite-state morphological analyzer (built with HFST) covering dialects from Gansu and Shaanxi, formalizing existing linguistic knowledge and evaluating its lexical coverage and ambiguity distribution across three text genres. The findings reveal that inflectional morphology in Dungan is rare—only 9.3% of tokens bear morphological markers—and 78.1% of lemmas are unambiguous, indicating a highly closed grammatical core with openness primarily at the lexical level. Out-of-vocabulary items account for 78%–95% of analysis failures, and integrating the morphological model improves lemma coverage by 5.2 percentage points.

0 citationsRead paper

First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers

Jul 18, 2026

This study investigates the tension between the predictability of unidirectional perturbations and the fragility of multidirectional combinations in local task adaptation of Transformers. Leveraging multi-task LoRA operating points, the authors systematically evaluate eight local adaptation properties across nine models of varying scales, establishing—via a preregistered protocol—the quantitative boundary between first-order predictability and pairwise combinatorial fragility for the first time. Integrating LoRA fine-tuning, task arithmetic, activation steering, gradient analysis, and Lie bracket modeling, they demonstrate that unidirectional perturbations are highly predictable within a first-order regime (effective window up to 10⁻²), yet over one-third of model–task pairs exhibit order sensitivity at even smaller scales. The Lie bracket precisely captures this ordering dependence, with a median ratio of predicted to empirical deviation of merely 1.002.

0 citationsRead paper

Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration

Jul 18, 2026

This work addresses the limitation of existing large model memory evaluations, which predominantly rely on dialogue or summarization and fail to assess scientific agents’ ability to recover evidence from full research papers. The authors propose the first full-text–oriented scientific memory evaluation paradigm, introducing two new benchmarks—PAIM and PTr—that formalize scientific memory as a budget-constrained, modality-aware context recovery task. Their approach combines sparse (BM25) and dense retrieval in a hybrid strategy, incorporating multi-granularity text ingestion, preservation of original content, controlled retrieval budgets, and consistency calibration across multiple evaluators, including humans. Experiments reveal that under strict retrieval budget constraints, the performance advantage of leading systems vanishes; hybrid retrieval substantially improves results; and LLM-based evaluators align closely with human judgments, achieving a resolution of 0.1 points.

0 citationsRead paper

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

Jul 07, 2026

Existing evaluations of code agents predominantly focus on task success or failure, overlooking the nuanced interaction trajectories observed in real-world usage. This work proposes the first fine-grained evaluation framework that integrates formal verification with large language model–generated commentary, leveraging complete interaction traces from actual production environments to establish a multidimensional and interpretable assessment mechanism. The framework enables detailed model behavior diagnostics, supports regression detection across model versions, and incorporates an automated evaluation pipeline. Its effectiveness has been validated through internal agent iterations and nightly regression testing, and the associated benchmark suite has been open-sourced to provide the research community with a high-value evaluation tool.

0 citationsRead paper

Recoverable but Not Stationary:Local Linear Structures in Weights and Activations

Jun 09, 2026

This study investigates the genuine existence and scale of local linear structures in large language models, critically examining the "fixed task plane" hypothesis. By analyzing weight and activation dynamics in synthetic multi-task Transformers and LoRA fine-tuning, the authors find that task gradients exhibit strong local low-rankness yet are not static. Building on this observation, they propose trajectory prefix bases to effectively capture recovery directions and develop a theoretical framework for local linearity under high-dimensional random search. Experiments on DistilGPT-2, GPT-2, and Qwen-0.5B demonstrate that local linear structures account for 77% of LoRA recovery displacement, and the cosine similarity between single-step gradients and CAA steering vectors reaches 0.58, providing significant empirical support for the interpretability of local linear structures in model behavior.

0 citationsRead paper