Bulbul: A Dataset for Dialectal Arabic Speech Recognition
为解决阿拉伯语自动语音识别面临的独特挑战,通过构建包含275名来自11个阿拉伯国家的说话者的多方言数据集BULBUL,并采用两层人工验证确保质量。
为解决阿拉伯语自动语音识别面临的独特挑战,通过构建包含275名来自11个阿拉伯国家的说话者的多方言数据集BULBUL,并采用两层人工验证确保质量。
This paper identifies a fundamental topological distinction between the solution spaces of 2-SAT and 3-SAT: while the 2-SAT solution space is contractible (admitting no nontrivial holes), the 3-SAT solution space contains exponentially many two-dimensional holes—formally, its second Betti number (b_2) admits an exponential lower bound. Method: We model Boolean formulas as cubical complexes embedded in the hypercube, analyze their homology via Betti numbers, construct explicit SAT reductions, and establish lower bounds within restricted query models. Contribution/Results: We introduce Betti numbers as paradigm-independent computational hardness invariants—demonstrating that high-dimensional topological obstructions constitute intrinsic barriers to efficiently solving 3-SAT. Crucially, these obstructions evade major complexity-theoretic barriers (relativization, natural proofs, algebrization). We rigorously prove an exponential lower bound on (b_2) for 3-SAT instances and derive exponential-time lower bounds for 3-SAT across multiple algorithmic models, providing structural evidence for (mathbf{P} eq mathbf{NP}).
Research on Arabic cyberbullying detection remains scarce, hindered by the scarcity of annotated data and inherent challenges in modeling low-resource languages. Method: This work introduces a high-quality, manually annotated dataset comprising 10,662 social media posts and proposes a hybrid architecture integrating BERT with LSTM/Bi-LSTM to enhance semantic representation of Arabic text. Annotation quality is rigorously validated via Cohen’s Kappa. Multiple ablation studies compare FastText embeddings, pre-trained BERT, and their combinations. Contribution/Results: The Bi-LSTM-BERT model achieves 97% accuracy, while Bi-LSTM augmented with FastText attains 98%, significantly outperforming baseline models. Notably, both variants demonstrate strong cross-domain generalization. This study establishes a reproducible, benchmark-quality dataset and an effective, transferable modeling framework for cyberbullying detection in low-resource languages—particularly Arabic.
This study addresses the binary classification of dental service providers—standard providers versus safety-net clinics (SNCs)—using real-world Medicare claims data (n = 24,300) with 38.1% missing values. To ensure robustness, we systematically compare three feature importance metrics—information gain, Gini impurity, and ANOVA—and identify volume of treatment services as the strongest discriminative feature. We empirically validate the resilience of stochastic gradient descent (SGD) and other algorithms to high-missingness data. Employing 12 machine learning algorithms with 10-fold cross-validation, neural networks achieve the highest accuracy (94.1%), followed closely by gradient boosting (93.2%) and random forests (93.0%). Ablation studies confirm that performance improves monotonically with the inclusion of key features. The work establishes an interpretable, highly robust modeling paradigm for healthcare resource categorization and safety-net service identification.
Large language models (LLMs) exhibit insufficient cultural awareness in Saudi Arabia’s linguistically diverse dialectal landscape and rich cultural context. Method: We introduce the first fine-grained, Saudi-specific cultural competence benchmark—covering five geographic regions and six cultural domains (e.g., cuisine, attire, festivals), incorporating open-ended, single-choice, and multiple-answer question formats, and distinguishing between commonsense and domain-specialized knowledge. We propose a novel “geographic–cultural two-dimensional decoupled evaluation framework” to isolate regional expertise. Contribution/Results: Evaluation across five state-of-the-art models—including GPT-4 and Llama 3.3—reveals a >37% average accuracy drop on region-specific questions and a 62% error rate on multiple-answer items, exposing critical deficits in localized cultural reasoning. The benchmark is constructed via expert annotation and cross-model consistency verification, advancing cultural assessment from generic to locale-specific evaluation and providing an essential empirical foundation for culturally adaptive LLM training.
为解决阿拉伯语自动语音识别面临的独特挑战,通过构建包含275名来自11个阿拉伯国家的说话者的多方言数据集BULBUL,并采用两层人工验证确保质量。
This paper identifies a fundamental topological distinction between the solution spaces of 2-SAT and 3-SAT: while the 2-SAT solution space is contractible (admitting no nontrivial holes), the 3-SAT solution space contains exponentially many two-dimensional holes—formally, its second Betti number (b_2) admits an exponential lower bound. Method: We model Boolean formulas as cubical complexes embedded in the hypercube, analyze their homology via Betti numbers, construct explicit SAT reductions, and establish lower bounds within restricted query models. Contribution/Results: We introduce Betti numbers as paradigm-independent computational hardness invariants—demonstrating that high-dimensional topological obstructions constitute intrinsic barriers to efficiently solving 3-SAT. Crucially, these obstructions evade major complexity-theoretic barriers (relativization, natural proofs, algebrization). We rigorously prove an exponential lower bound on (b_2) for 3-SAT instances and derive exponential-time lower bounds for 3-SAT across multiple algorithmic models, providing structural evidence for (mathbf{P} eq mathbf{NP}).
Research on Arabic cyberbullying detection remains scarce, hindered by the scarcity of annotated data and inherent challenges in modeling low-resource languages. Method: This work introduces a high-quality, manually annotated dataset comprising 10,662 social media posts and proposes a hybrid architecture integrating BERT with LSTM/Bi-LSTM to enhance semantic representation of Arabic text. Annotation quality is rigorously validated via Cohen’s Kappa. Multiple ablation studies compare FastText embeddings, pre-trained BERT, and their combinations. Contribution/Results: The Bi-LSTM-BERT model achieves 97% accuracy, while Bi-LSTM augmented with FastText attains 98%, significantly outperforming baseline models. Notably, both variants demonstrate strong cross-domain generalization. This study establishes a reproducible, benchmark-quality dataset and an effective, transferable modeling framework for cyberbullying detection in low-resource languages—particularly Arabic.
This study addresses the binary classification of dental service providers—standard providers versus safety-net clinics (SNCs)—using real-world Medicare claims data (n = 24,300) with 38.1% missing values. To ensure robustness, we systematically compare three feature importance metrics—information gain, Gini impurity, and ANOVA—and identify volume of treatment services as the strongest discriminative feature. We empirically validate the resilience of stochastic gradient descent (SGD) and other algorithms to high-missingness data. Employing 12 machine learning algorithms with 10-fold cross-validation, neural networks achieve the highest accuracy (94.1%), followed closely by gradient boosting (93.2%) and random forests (93.0%). Ablation studies confirm that performance improves monotonically with the inclusion of key features. The work establishes an interpretable, highly robust modeling paradigm for healthcare resource categorization and safety-net service identification.
Large language models (LLMs) exhibit insufficient cultural awareness in Saudi Arabia’s linguistically diverse dialectal landscape and rich cultural context. Method: We introduce the first fine-grained, Saudi-specific cultural competence benchmark—covering five geographic regions and six cultural domains (e.g., cuisine, attire, festivals), incorporating open-ended, single-choice, and multiple-answer question formats, and distinguishing between commonsense and domain-specialized knowledge. We propose a novel “geographic–cultural two-dimensional decoupled evaluation framework” to isolate regional expertise. Contribution/Results: Evaluation across five state-of-the-art models—including GPT-4 and Llama 3.3—reveals a >37% average accuracy drop on region-specific questions and a 62% error rate on multiple-answer items, exposing critical deficits in localized cultural reasoning. The benchmark is constructed via expert annotation and cross-model consistency verification, advancing cultural assessment from generic to locale-specific evaluation and providing an essential empirical foundation for culturally adaptive LLM training.