Institution profile

Kyoto University of Advanced Science

Academic institutionasia · jp
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Conceptual Design of an Ecosystem for Real Farm Data Collection toward Agricultural AI Foundation Models

Jun 22, 2026

This study addresses the scarcity of real-world farm data, insufficient incentives for data contributors, and data authenticity challenges exacerbated by generative AI, which collectively hinder the development of agricultural foundation models. To overcome these limitations, this work proposes a sustainable data collection and distribution ecosystem that uniquely integrates economic incentives, authenticity verification, and AI data requirements. The system features an automated pricing mechanism driven by data demand and rarity, a revenue-sharing strategy tailored for farmers, and certified device uploads to ensure data trustworthiness, complemented by an economic value estimation model. Empirical results demonstrate that the proposed framework is economically sustainable and effectively achieves a tripartite win-win among farmers, AI enterprises, and the platform, thereby providing a high-quality, continuous stream of authentic data essential for agricultural robotics and foundation models.

0 citationsRead paper

Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources

Jul 12, 2025

To address the scarcity of high-quality romanization resources for Sinhala, this paper proposes the first systematic transliteration framework—hybridizing sequence-to-sequence modeling, phoneme-level rule-based mapping, and post-processing optimization. We introduce the first standardized, open-source Sinhala romanization–native script parallel corpus (120K+ high-quality pairs), constructed by unifying heterogeneous transcription data from multiple sources. We further release a reusable toolchain and a benchmark evaluation set. Experiments demonstrate substantial improvements over existing tools: +8.3 BLEU points and +12.7% character-level F1 score. This work fills a critical technical gap in transliteration modeling for low-resource South Asian languages and establishes foundational support for downstream Sinhala NLP tasks, including automatic speech recognition (ASR) and machine translation.

0 citationsRead paper

Integrated ensemble of BERT- and features-based models for authorship attribution in Japanese literary works

Apr 11, 2025

This paper addresses the few-shot Japanese literary author attribution (AA) task. Methodologically, it proposes a dual-path ensemble framework integrating traditional stylistic features with pre-trained language models (PLMs). It is the first to empirically validate BERT’s effectiveness for few-shot Japanese AA; simultaneously, TF-IDF and character n-gram features are extracted and modeled using multi-layer classifiers—XGBoost, SVM, and MLP—with weighted voting for fusion. The core contribution lies in establishing a synergistic “PLM + traditional features” dual-path paradigm, substantially improving generalization under data-scarce conditions. Experiments on test sets disjoint from pre-training data demonstrate an approximately 14-percentage-point improvement in macro-F1 score over baseline models. The integrated model consistently outperforms all individual components, establishing a new state-of-the-art performance ceiling for few-shot Japanese AA.

0 citationsRead paper
Recent publications

Latest Papers

Conceptual Design of an Ecosystem for Real Farm Data Collection toward Agricultural AI Foundation Models

Jun 22, 2026

This study addresses the scarcity of real-world farm data, insufficient incentives for data contributors, and data authenticity challenges exacerbated by generative AI, which collectively hinder the development of agricultural foundation models. To overcome these limitations, this work proposes a sustainable data collection and distribution ecosystem that uniquely integrates economic incentives, authenticity verification, and AI data requirements. The system features an automated pricing mechanism driven by data demand and rarity, a revenue-sharing strategy tailored for farmers, and certified device uploads to ensure data trustworthiness, complemented by an economic value estimation model. Empirical results demonstrate that the proposed framework is economically sustainable and effectively achieves a tripartite win-win among farmers, AI enterprises, and the platform, thereby providing a high-quality, continuous stream of authentic data essential for agricultural robotics and foundation models.

0 citationsRead paper

Swa-bhasha Resource Hub: Romanized Sinhala to Sinhala Transliteration Systems and Data Resources

Jul 12, 2025

To address the scarcity of high-quality romanization resources for Sinhala, this paper proposes the first systematic transliteration framework—hybridizing sequence-to-sequence modeling, phoneme-level rule-based mapping, and post-processing optimization. We introduce the first standardized, open-source Sinhala romanization–native script parallel corpus (120K+ high-quality pairs), constructed by unifying heterogeneous transcription data from multiple sources. We further release a reusable toolchain and a benchmark evaluation set. Experiments demonstrate substantial improvements over existing tools: +8.3 BLEU points and +12.7% character-level F1 score. This work fills a critical technical gap in transliteration modeling for low-resource South Asian languages and establishes foundational support for downstream Sinhala NLP tasks, including automatic speech recognition (ASR) and machine translation.

0 citationsRead paper

Integrated ensemble of BERT- and features-based models for authorship attribution in Japanese literary works

Apr 11, 2025

This paper addresses the few-shot Japanese literary author attribution (AA) task. Methodologically, it proposes a dual-path ensemble framework integrating traditional stylistic features with pre-trained language models (PLMs). It is the first to empirically validate BERT’s effectiveness for few-shot Japanese AA; simultaneously, TF-IDF and character n-gram features are extracted and modeled using multi-layer classifiers—XGBoost, SVM, and MLP—with weighted voting for fusion. The core contribution lies in establishing a synergistic “PLM + traditional features” dual-path paradigm, substantially improving generalization under data-scarce conditions. Experiments on test sets disjoint from pre-training data demonstrate an approximately 14-percentage-point improvement in macro-F1 score over baseline models. The integrated model consistently outperforms all individual components, establishing a new state-of-the-art performance ceiling for few-shot Japanese AA.

0 citationsRead paper