Institution profile

National Research and Innovation Agency

Academic institutionasia · id
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Linking Hadith Narrator Identities Across Heterogeneous Arabic Biographical Databases: A Multi-Signal Entity Resolution Pipeline

Jun 30, 2026

This study addresses the challenge of identity linkage among hundreds of thousands of Islamic hadith transmitters across heterogeneous Arabic biographical databases, where the absence of unified identifiers impedes cross-resource integration. To resolve this, the work proposes the first cross-database, multi-signal entity resolution framework tailored for Arabic hadith transmitters. The approach employs a two-stage pipeline: first linking transmitters from the Sanadset corpus to the HadithTransmitters database via name similarity, then aligning with the MuslimScholars database through a weighted fusion of multiple signals—namely name, death year, and reliability rating—augmented by a transitive linking strategy. Operating without metadata, the framework achieves high-coverage identity resolution, constructing a directed transmission graph comprising 185,216 nodes and 814,093 edges. This effort yields the first structured integration of these three major resources, accompanied by the public release of high-quality linked corpora and a cross-source biographical knowledge graph.

0 citationsRead paper

IndoBERT-Sentiment: Context-Conditioned Sentiment Classification for Indonesian Text

Apr 08, 2026

This study addresses the challenge of context-dependent sentiment misclassification in Indonesian text, a limitation commonly observed in existing models that disregard topical context. To overcome this, the authors propose a joint modeling approach that integrates topical context with textual content, marking the first successful application of context-conditioned modeling to Indonesian sentiment analysis. Built upon the IndoBERT Large architecture (335 million parameters), the model is trained on 31,360 context–text pairs spanning 188 distinct topics. Evaluated on a held-out test set, it achieves an accuracy of 88.1% and a macro F1-score of 0.856, representing a substantial improvement of 35.6 F1 points over the strongest baseline. This advancement demonstrates the critical role of contextual information in enhancing sentiment classification performance for Indonesian.

0 citationsRead paper

IndoBERT-Relevancy: A Context-Conditioned Relevancy Classifier for Indonesian Text

Mar 27, 2026

This work addresses the lack of effective models for Indonesian relevance classification by proposing a context-conditioned relevance classifier built upon IndoBERT Large (335 million parameters). Leveraging an iterative, failure-driven data construction methodology, the authors curate a novel dataset comprising 31,360 annotated sentence pairs spanning 188 topics, which integrates both authentic multi-source data and targeted synthetic examples to encompass both formal and informal Indonesian text. The resulting model achieves 96.5% accuracy and an F1 score of 0.948 on the test set, demonstrating substantially improved robustness. The model and dataset have been publicly released on the Hugging Face platform to support further research in Indonesian natural language processing.

0 citationsRead paper

SocialX: A Modular Platform for Multi-Source Big Data Research in Indonesia

Mar 27, 2026

This study addresses the persistent challenge of data fragmentation in Indonesian multi-source big data research, where heterogeneous data formats, access protocols, and noise characteristics across platforms often necessitate redundant development of collection and analysis pipelines. To overcome this, the authors propose a modular, source-agnostic unified processing platform featuring a three-layer decoupled architecture—comprising data ingestion, language-aware preprocessing, and pluggable analytics—alongside a lightweight task orchestration mechanism. This design enables seamless integration of new data sources, preprocessing methods, or analytical tools. The platform, publicly available at https://www.socialx.id, has been validated through representative workflows, demonstrating a significant reduction in the barrier to entry for multi-source big data research while enhancing scalability and reusability.

0 citationsRead paper

Mobile Robot Localization via Indoor Positioning System and Odometry Fusion

Sep 20, 2025

Indoor mobile robot localization suffers from multipath interference in standalone ultrasonic indoor positioning systems (IPS) and cumulative drift in wheel odometry. To address these limitations, this paper proposes a tightly coupled multi-sensor fusion method based on the extended Kalman filter (EKF), jointly estimating robot pose using IPS and wheel odometry. The approach leverages their complementary strengths: IPS provides global reference corrections, while odometry delivers high-frequency motion continuity; the EKF explicitly models and mitigates nonlinear uncertainties—including wheel slip and measurement noise. Experimental results demonstrate that the proposed fusion scheme reduces average localization error by approximately 62%, significantly suppresses trajectory drift, and markedly improves robustness and long-term stability compared to single-sensor baselines. This solution enables high-accuracy, infrastructure-light indoor autonomous navigation at low hardware cost.

0 citationsRead paper
Recent publications

Latest Papers

Linking Hadith Narrator Identities Across Heterogeneous Arabic Biographical Databases: A Multi-Signal Entity Resolution Pipeline

Jun 30, 2026

This study addresses the challenge of identity linkage among hundreds of thousands of Islamic hadith transmitters across heterogeneous Arabic biographical databases, where the absence of unified identifiers impedes cross-resource integration. To resolve this, the work proposes the first cross-database, multi-signal entity resolution framework tailored for Arabic hadith transmitters. The approach employs a two-stage pipeline: first linking transmitters from the Sanadset corpus to the HadithTransmitters database via name similarity, then aligning with the MuslimScholars database through a weighted fusion of multiple signals—namely name, death year, and reliability rating—augmented by a transitive linking strategy. Operating without metadata, the framework achieves high-coverage identity resolution, constructing a directed transmission graph comprising 185,216 nodes and 814,093 edges. This effort yields the first structured integration of these three major resources, accompanied by the public release of high-quality linked corpora and a cross-source biographical knowledge graph.

0 citationsRead paper

IndoBERT-Sentiment: Context-Conditioned Sentiment Classification for Indonesian Text

Apr 08, 2026

This study addresses the challenge of context-dependent sentiment misclassification in Indonesian text, a limitation commonly observed in existing models that disregard topical context. To overcome this, the authors propose a joint modeling approach that integrates topical context with textual content, marking the first successful application of context-conditioned modeling to Indonesian sentiment analysis. Built upon the IndoBERT Large architecture (335 million parameters), the model is trained on 31,360 context–text pairs spanning 188 distinct topics. Evaluated on a held-out test set, it achieves an accuracy of 88.1% and a macro F1-score of 0.856, representing a substantial improvement of 35.6 F1 points over the strongest baseline. This advancement demonstrates the critical role of contextual information in enhancing sentiment classification performance for Indonesian.

0 citationsRead paper

IndoBERT-Relevancy: A Context-Conditioned Relevancy Classifier for Indonesian Text

Mar 27, 2026

This work addresses the lack of effective models for Indonesian relevance classification by proposing a context-conditioned relevance classifier built upon IndoBERT Large (335 million parameters). Leveraging an iterative, failure-driven data construction methodology, the authors curate a novel dataset comprising 31,360 annotated sentence pairs spanning 188 topics, which integrates both authentic multi-source data and targeted synthetic examples to encompass both formal and informal Indonesian text. The resulting model achieves 96.5% accuracy and an F1 score of 0.948 on the test set, demonstrating substantially improved robustness. The model and dataset have been publicly released on the Hugging Face platform to support further research in Indonesian natural language processing.

0 citationsRead paper

SocialX: A Modular Platform for Multi-Source Big Data Research in Indonesia

Mar 27, 2026

This study addresses the persistent challenge of data fragmentation in Indonesian multi-source big data research, where heterogeneous data formats, access protocols, and noise characteristics across platforms often necessitate redundant development of collection and analysis pipelines. To overcome this, the authors propose a modular, source-agnostic unified processing platform featuring a three-layer decoupled architecture—comprising data ingestion, language-aware preprocessing, and pluggable analytics—alongside a lightweight task orchestration mechanism. This design enables seamless integration of new data sources, preprocessing methods, or analytical tools. The platform, publicly available at https://www.socialx.id, has been validated through representative workflows, demonstrating a significant reduction in the barrier to entry for multi-source big data research while enhancing scalability and reusability.

0 citationsRead paper

Mobile Robot Localization via Indoor Positioning System and Odometry Fusion

Sep 20, 2025

Indoor mobile robot localization suffers from multipath interference in standalone ultrasonic indoor positioning systems (IPS) and cumulative drift in wheel odometry. To address these limitations, this paper proposes a tightly coupled multi-sensor fusion method based on the extended Kalman filter (EKF), jointly estimating robot pose using IPS and wheel odometry. The approach leverages their complementary strengths: IPS provides global reference corrections, while odometry delivers high-frequency motion continuity; the EKF explicitly models and mitigates nonlinear uncertainties—including wheel slip and measurement noise. Experimental results demonstrate that the proposed fusion scheme reduces average localization error by approximately 62%, significantly suppresses trajectory drift, and markedly improves robustness and long-term stability compared to single-sensor baselines. This solution enables high-accuracy, infrastructure-light indoor autonomous navigation at low hardware cost.

0 citationsRead paper