Institution profile

SberDevices

Industry researcheurope · ru
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Instruction-Following Evaluation in Function Calling for Large Language Models

Sep 22, 2025

Existing function-calling evaluation benchmarks (e.g., BFCL, tau²-Bench) assess only parameter semantic correctness, neglecting models’ adherence to syntactic formatting instructions—such as quoted strings or ISO date formats. Method: We introduce the first benchmark explicitly targeting *format instruction adherence*, comprising 750 test cases with explicit, diverse formatting constraints. Leveraging embedded JSON Schema validation rules, our benchmark enables fully automated, algorithmic evaluation. Contribution/Results: It is the first to systematically quantify large language models’ precision in executing embedded formatting directives within structured outputs. Experiments reveal alarmingly high error rates—30%–60%—even for state-of-the-art models (e.g., GPT-5, Claude 4.1 Opus) on basic formatting requirements, exposing a critical reliability gap for AI agents in real-world deployment.

0 citationsRead paper

HandReader: Advanced Techniques for Efficient Fingerspelling Recognition

May 15, 2025

This work addresses three key challenges in sign language fingerspelling recognition: difficulty modeling rapid hand motion, weak handling of variable-length video sequences, and low accuracy in recognizing proper nouns. To tackle these, we propose: (1) a Temporal Shift Adaptive Module (TSAM) and a Temporal Pose Encoder (TPE) for joint temporal modeling of RGB and keypoint features; (2) the first RGB-keypoint multimodal fusion architecture, integrating 2D/3D convolutions, temporal feature alignment, and keypoint tensor representation; and (3) Znaki—the first publicly available Russian fingerspelling dataset, which we construct and open-source. Our method achieves state-of-the-art performance on ChicagoFSWild, ChicagoFSWild+, and Znaki. Both the model and the Znaki dataset are publicly released to foster further research.

0 citationsRead paper

RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

Feb 11, 2025

Text-to-image generation models exhibit pervasive English cultural bias, leading to misrepresentation and stereotyping of non-English cultures—such as Russian culture—and thereby introducing risks of bias and offense. To address this, we introduce RusCode, the first benchmark for evaluating visual representation of Russian culture. It systematically defines 19 Russian cultural visual codes and comprises a bilingual (Russian–English) prompt set of 1,250 instances. Our methodology integrates culturally sensitive prompt engineering, a human-annotated taxonomy, and multidimensional human evaluation—assessing faithfulness, recognizability, and cultural appropriateness. Experiments reveal substantial inaccuracies in mainstream models’ generation of Russian cultural concepts. This work fills a critical gap in cross-cultural generative evaluation within computer vision, establishing a reproducible diagnostic baseline and a culturally adaptive evaluation paradigm.

0 citationsRead paper
Recent publications

Latest Papers

Instruction-Following Evaluation in Function Calling for Large Language Models

Sep 22, 2025

Existing function-calling evaluation benchmarks (e.g., BFCL, tau²-Bench) assess only parameter semantic correctness, neglecting models’ adherence to syntactic formatting instructions—such as quoted strings or ISO date formats. Method: We introduce the first benchmark explicitly targeting *format instruction adherence*, comprising 750 test cases with explicit, diverse formatting constraints. Leveraging embedded JSON Schema validation rules, our benchmark enables fully automated, algorithmic evaluation. Contribution/Results: It is the first to systematically quantify large language models’ precision in executing embedded formatting directives within structured outputs. Experiments reveal alarmingly high error rates—30%–60%—even for state-of-the-art models (e.g., GPT-5, Claude 4.1 Opus) on basic formatting requirements, exposing a critical reliability gap for AI agents in real-world deployment.

0 citationsRead paper

HandReader: Advanced Techniques for Efficient Fingerspelling Recognition

May 15, 2025

This work addresses three key challenges in sign language fingerspelling recognition: difficulty modeling rapid hand motion, weak handling of variable-length video sequences, and low accuracy in recognizing proper nouns. To tackle these, we propose: (1) a Temporal Shift Adaptive Module (TSAM) and a Temporal Pose Encoder (TPE) for joint temporal modeling of RGB and keypoint features; (2) the first RGB-keypoint multimodal fusion architecture, integrating 2D/3D convolutions, temporal feature alignment, and keypoint tensor representation; and (3) Znaki—the first publicly available Russian fingerspelling dataset, which we construct and open-source. Our method achieves state-of-the-art performance on ChicagoFSWild, ChicagoFSWild+, and Znaki. Both the model and the Znaki dataset are publicly released to foster further research.

0 citationsRead paper

RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

Feb 11, 2025

Text-to-image generation models exhibit pervasive English cultural bias, leading to misrepresentation and stereotyping of non-English cultures—such as Russian culture—and thereby introducing risks of bias and offense. To address this, we introduce RusCode, the first benchmark for evaluating visual representation of Russian culture. It systematically defines 19 Russian cultural visual codes and comprises a bilingual (Russian–English) prompt set of 1,250 instances. Our methodology integrates culturally sensitive prompt engineering, a human-annotated taxonomy, and multidimensional human evaluation—assessing faithfulness, recognizability, and cultural appropriateness. Experiments reveal substantial inaccuracies in mainstream models’ generation of Russian cultural concepts. This work fills a critical gap in cross-cultural generative evaluation within computer vision, establishing a reproducible diagnostic baseline and a culturally adaptive evaluation paradigm.

0 citationsRead paper