Institution profile

National Technical University (Kiev Politechnical Institute)

Academic institutioneurope · ua
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding

May 11, 2026

This work addresses the challenge of automatically answering multiple-choice questions from Ukrainian multi-domain PDF documents while simultaneously locating supporting evidence. The authors propose a retrieval-augmented question-answering pipeline that integrates context-aware PDF chunking, dense retrieval informed by both questions and answer options, and a reranking stage enhanced with awareness of the answer space. A constrained decoding mechanism is introduced during answer generation to ensure correctness under strict competition constraints, preserving document structure without relying on complex post-processing. The system leverages Qwen3-Embedding-8B for retrieval, a fine-tuned Qwen3-Reranker-8B for reranking, and Qwen3-32B for final answer selection. Reranking improves Recall@1 from 0.6957 to 0.7935, and using the top two retrieved passages boosts accuracy from 0.9348 to 0.9674, achieving a private leaderboard score of 0.9598.

0 citationsRead paper

КЛАСИФІКАЦІЯ ЛЕГЕНЕВИХ УТВОРІВ НА ФРАГМЕНТАХ КТ-ЗОБРАЖЕНЬ ІЗ ВИКОРИСТАННЯМ ТРИВИМІРНИХ ЗГОРТКОВИХ НЕЙРОННИХ МЕРЕЖ

Dec 30, 2025Таврійський науковий вісник Серія Технічні науки

Рак легенів залишається одним із найпоширеніших і найсмертельніших видів раку у світі. Ймовірність успішного лікування значною мірою залежить від стадії, на якій було встановлено діагноз. Тому раннє виявлення раку легенів є надзвичайно важливим медичним завданням. Проте ця задача створює значні труднощі для торакальних радіологів через велику кількість досліджень, які необхідно проаналізувати, наявність множинних утворів у легенях і невеликі розміри багатьох із них, що ускладнює візуальну оцінку. У зв’язку з цим розробка автоматизованих систем, які містять високоточні та обчислювально ефективні модулі виявлення й класифікації легеневих утворів, є вкрай актуальною. У цьому дослідженні представлено три методологічні вдосконалення для задачі класифікації легеневих утворів: (1) удосконалену стратегію фрагментування КТ-зображень, яка дозволяє моделі зосередитися на цільовому утворі та зменшити обчислювальні витрати; (2) методи фільтрації цільових міток для усунення зашумлених анотацій; (3) нові методи аугментації, спрямовані на підвищення стійкості моделі. Інтеграція цих підходів дозволяє створити надійну підсистему класифікації у складі комплексної системи підтримки прийняття клінічних рішень для виявлення раку легенів, здатної працювати з різними протоколами проведення КТ дослідження, типами сканерів і вхідними моделями (сегментації чи детекції). Розглянуто дві постановки задачі – багатокласову та бінарну, а також кілька методів агрегації для перетворення багатокласових прогнозів у бінарні. Проведено розширене дослідження для кількісної оцінки внеску кожного запропонованого методологічного вдосконалення. Багатокласова модель досягла Macro ROC AUC = 0.9176 і Macro F1 = 0.7658, тоді як бінарна модель показала Binary ROC AUC = 0.9383 і Binary F1 = 0.8668 на наборі даних LIDC-IDRI. Отримані результати перевищують показники низки попередніх підходів і демонструють ефективність кращу від сучасного стану галузі для цієї задачі.

0 citationsRead paper

Vuyko Mistral: Adapting LLMs for Low-Resource Dialectal Translation

Jun 09, 2025

This study addresses the critical challenge of low-resource, morphologically complex Hutsul dialect—a regional variety of Ukrainian. We present the first Hutsul–Standard Ukrainian parallel corpus and lexicon. To overcome data scarcity, we propose a RAG-enhanced synthetic data generation pipeline and fine-tune an open-source 7B LLM using LoRA. We introduce the first LLM adaptation framework specifically designed for low-resource dialects and establish an unsupervised, multi-metric evaluation suite integrating BLEU, chrF++, TER, and GPT-4o-based discriminative assessment. Experimental results demonstrate that our fine-tuned model significantly outperforms GPT-4o in zero-shot translation across all metrics. All resources—including the corpus, dictionary, trained models, and code—are publicly released, constituting foundational infrastructure and a methodological blueprint for low-resource dialect NLP research.

0 citationsRead paper

Cerny type automata and rank conjecture

Jan 31, 2025

This work addresses the unified verification of the Černý conjecture and the rank conjecture for Černý-type synchronizing automata and transformation monoids generated by simple idempotents and regular permutation groups. Method: We introduce, for the first time, a unifying structural framework—termed the Černý-type framework—that integrates combinatorial semigroup theory, idempotent decomposition, orbit analysis of permutation groups, and synchronizing automata techniques. Contribution/Results: We rigorously prove that such automata satisfy the Černý conjecture and their associated monoids satisfy the rank conjecture. Moreover, we derive a tight upper bound on the reset threshold—namely, the optimal bound—thereby transcending traditional analyses confined to isolated automaton models. This constitutes the first theoretical characterization of worst-case reset behavior in synchronizing systems that is both universally applicable across this broad class and quantitatively precise.

0 citationsRead paper
Recent publications

Latest Papers

Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding

May 11, 2026

This work addresses the challenge of automatically answering multiple-choice questions from Ukrainian multi-domain PDF documents while simultaneously locating supporting evidence. The authors propose a retrieval-augmented question-answering pipeline that integrates context-aware PDF chunking, dense retrieval informed by both questions and answer options, and a reranking stage enhanced with awareness of the answer space. A constrained decoding mechanism is introduced during answer generation to ensure correctness under strict competition constraints, preserving document structure without relying on complex post-processing. The system leverages Qwen3-Embedding-8B for retrieval, a fine-tuned Qwen3-Reranker-8B for reranking, and Qwen3-32B for final answer selection. Reranking improves Recall@1 from 0.6957 to 0.7935, and using the top two retrieved passages boosts accuracy from 0.9348 to 0.9674, achieving a private leaderboard score of 0.9598.

0 citationsRead paper

КЛАСИФІКАЦІЯ ЛЕГЕНЕВИХ УТВОРІВ НА ФРАГМЕНТАХ КТ-ЗОБРАЖЕНЬ ІЗ ВИКОРИСТАННЯМ ТРИВИМІРНИХ ЗГОРТКОВИХ НЕЙРОННИХ МЕРЕЖ

Dec 30, 2025Таврійський науковий вісник Серія Технічні науки

Рак легенів залишається одним із найпоширеніших і найсмертельніших видів раку у світі. Ймовірність успішного лікування значною мірою залежить від стадії, на якій було встановлено діагноз. Тому раннє виявлення раку легенів є надзвичайно важливим медичним завданням. Проте ця задача створює значні труднощі для торакальних радіологів через велику кількість досліджень, які необхідно проаналізувати, наявність множинних утворів у легенях і невеликі розміри багатьох із них, що ускладнює візуальну оцінку. У зв’язку з цим розробка автоматизованих систем, які містять високоточні та обчислювально ефективні модулі виявлення й класифікації легеневих утворів, є вкрай актуальною. У цьому дослідженні представлено три методологічні вдосконалення для задачі класифікації легеневих утворів: (1) удосконалену стратегію фрагментування КТ-зображень, яка дозволяє моделі зосередитися на цільовому утворі та зменшити обчислювальні витрати; (2) методи фільтрації цільових міток для усунення зашумлених анотацій; (3) нові методи аугментації, спрямовані на підвищення стійкості моделі. Інтеграція цих підходів дозволяє створити надійну підсистему класифікації у складі комплексної системи підтримки прийняття клінічних рішень для виявлення раку легенів, здатної працювати з різними протоколами проведення КТ дослідження, типами сканерів і вхідними моделями (сегментації чи детекції). Розглянуто дві постановки задачі – багатокласову та бінарну, а також кілька методів агрегації для перетворення багатокласових прогнозів у бінарні. Проведено розширене дослідження для кількісної оцінки внеску кожного запропонованого методологічного вдосконалення. Багатокласова модель досягла Macro ROC AUC = 0.9176 і Macro F1 = 0.7658, тоді як бінарна модель показала Binary ROC AUC = 0.9383 і Binary F1 = 0.8668 на наборі даних LIDC-IDRI. Отримані результати перевищують показники низки попередніх підходів і демонструють ефективність кращу від сучасного стану галузі для цієї задачі.

0 citationsRead paper

Vuyko Mistral: Adapting LLMs for Low-Resource Dialectal Translation

Jun 09, 2025

This study addresses the critical challenge of low-resource, morphologically complex Hutsul dialect—a regional variety of Ukrainian. We present the first Hutsul–Standard Ukrainian parallel corpus and lexicon. To overcome data scarcity, we propose a RAG-enhanced synthetic data generation pipeline and fine-tune an open-source 7B LLM using LoRA. We introduce the first LLM adaptation framework specifically designed for low-resource dialects and establish an unsupervised, multi-metric evaluation suite integrating BLEU, chrF++, TER, and GPT-4o-based discriminative assessment. Experimental results demonstrate that our fine-tuned model significantly outperforms GPT-4o in zero-shot translation across all metrics. All resources—including the corpus, dictionary, trained models, and code—are publicly released, constituting foundational infrastructure and a methodological blueprint for low-resource dialect NLP research.

0 citationsRead paper

Cerny type automata and rank conjecture

Jan 31, 2025

This work addresses the unified verification of the Černý conjecture and the rank conjecture for Černý-type synchronizing automata and transformation monoids generated by simple idempotents and regular permutation groups. Method: We introduce, for the first time, a unifying structural framework—termed the Černý-type framework—that integrates combinatorial semigroup theory, idempotent decomposition, orbit analysis of permutation groups, and synchronizing automata techniques. Contribution/Results: We rigorously prove that such automata satisfy the Černý conjecture and their associated monoids satisfy the rank conjecture. Moreover, we derive a tight upper bound on the reset threshold—namely, the optimal bound—thereby transcending traditional analyses confined to isolated automaton models. This constitutes the first theoretical characterization of worst-case reset behavior in synchronizing systems that is both universally applicable across this broad class and quantitatively precise.

0 citationsRead paper