Q-Sat AI: Machine Learning-Based Decision Support for Data Saturation in Qualitative Studies

📅 2025-11-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Qualitative research often relies on subjective judgments of data saturation to determine sample size, leading to methodological inconsistency and compromised rigor. To address this, we propose the first machine learning–based decision support model for sample size determination, pioneering the application of ensemble learning methods—including XGBoost and Random Forest—to quantitatively model the data saturation process. The model integrates ten key study design parameters, undergoes rigorous preprocessing and outlier removal, and achieves an R² of 0.85, effectively capturing nonlinear relationships in sampling dynamics. Feature importance analysis empirically validates foundational theoretical assumptions—such as the influence of study type and informational power—advancing standardization in qualitative methodology. The model has been implemented as an open-source web tool for researchers and reviewers, substantially enhancing transparency, reproducibility, and methodological rigor in sample size justification.

Technology Category

Application Category

📝 Abstract
The determination of sample size in qualitative research has traditionally relied on the subjective and often ambiguous principle of data saturation, which can lead to inconsistencies and threaten methodological rigor. This study introduces a new, systematic model based on machine learning (ML) to make this process more objective. Utilizing a dataset derived from five fundamental qualitative research approaches - namely, Case Study, Grounded Theory, Phenomenology, Narrative Research, and Ethnographic Research - we developed an ensemble learning model. Ten critical parameters, including research scope, information power, and researcher competence, were evaluated using an ordinal scale and used as input features. After thorough preprocessing and outlier removal, multiple ML algorithms were trained and compared. The K-Nearest Neighbors (KNN), Gradient Boosting (GB), Random Forest (RF), XGBoost, and Decision Tree (DT) algorithms showed the highest explanatory power (Test R2 ~ 0.85), effectively modeling the complex, non-linear relationships involved in qualitative sampling decisions. Feature importance analysis confirmed the vital roles of research design type and information power, providing quantitative validation of key theoretical assumptions in qualitative methodology. The study concludes by proposing a conceptual framework for a web-based computational application designed to serve as a decision support system for qualitative researchers, journal reviewers, and thesis advisors. This model represents a significant step toward standardizing sample size justification, enhancing transparency, and strengthening the epistemological foundation of qualitative inquiry through evidence-based, systematic decision-making.
Problem

Research questions and friction points this paper is trying to address.

Objectively determining qualitative sample size using machine learning models
Addressing data saturation subjectivity through ensemble learning algorithms
Developing computational decision support for qualitative research standardization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Machine learning ensemble model for data saturation
Ten parameters evaluated using ordinal scale inputs
Web-based computational application for decision support
💼 Related Jobs
No related jobs found.
H
Hasan Tutar
Bolu Abant Izzet Baysal University, Faculty of Communication, 14030 – Merkez, Bolu, Turkey
C
Caner Erden
Department of Computer Engineering, Faculty of Technology, Sakarya University of Applied Science, Sakarya, Turkey
Ü
Ümit Şentürk
Department of Computer Engineering, Faculty of Engineering, Bolu Abant Izzet Baysal University, 14280, Bolu, Turkey