A Security Risk Assessment Framework for AI-Powered Development Tools
本文提出了一种安全风险评估框架(SRF),通过结合威胁建模、安全分析和基于漏洞关键性的定量风险评估方法,来解决AI生成代码的安全风险评估问题。
本文提出了一种安全风险评估框架(SRF),通过结合威胁建模、安全分析和基于漏洞关键性的定量风险评估方法,来解决AI生成代码的安全风险评估问题。
This study addresses depression and anxiety among university students stemming from career-related stress by proposing a privacy-preserving, culturally sensitive framework for early mental health risk detection. The approach integrates structured behavioral data with facial emotion cues extracted from interview videos, employing an attention-augmented intermediate fusion neural network for modeling. To enable collaborative training across institutions without sharing raw data, the framework leverages federated learning. Model interpretability is enhanced through Integrated Gradients and SHAP analyses, revealing key behavioral markers—such as gaze avoidance and reduced facial expressivity—that align with established psychological theories. Evaluated on a dataset of Pakistani university students, the model achieves 92.08% accuracy and an F1 score of 89.12%, demonstrating its effectiveness in identifying early signs of psychological distress.
This study systematically evaluates the performance of large language models on Arabic numeral reading tasks—including both Eastern and Western Arabic-Indic digits—across six contextual categories: plain numbers, addresses, dates, quantities, prices, and others. Using a comprehensive benchmark comprising 210 tasks and 59,010 test cases, the authors assess 71 models under four prompting strategies: zero-shot, zero-shot chain-of-thought (CoT), few-shot, and few-shot CoT. For the first time, numerical accuracy is disentangled from instruction-following capability, revealing a wide performance range (14.29%–99.05%) and demonstrating that few-shot CoT improves accuracy by 2.8× over zero-shot. Notably, only six models consistently produce structured outputs, highlighting that high accuracy does not guarantee reliable structured responses; thus, post-processing remains essential for practical deployment.
This study presents the first systematic evaluation of gender and cultural representation biases in mainstream text-to-image models (ImageFX, DALL-E 3, Grok) when generating occupational imagery for Saudi contexts. Using neutral occupation prompts, we generated 1,006 images; two Saudi native annotators performed double-blind labeling on gender, attire, setting, activity, and age, with discrepancies resolved by a senior researcher and supplemented by LLM-assisted analysis. Results reveal pervasive male bias—e.g., 96% male representation in DALL-E 3 outputs—as well as over-masculinization of technical and leadership roles and recurrent cultural misrepresentations in attire and environmental cues. We attribute these biases primarily to cultural homogeneity and societal stereotypes embedded in training data, rather than inherent model limitations. The work establishes the first Arabic-context-specific benchmark and methodological framework for assessing occupational representation fairness in generative AI, advancing equitable AI evaluation practices.
Arabic, the fourth most widely used language online, suffers from a longstanding scarcity of multi-task annotated resources for propaganda, sentiment, and emotion analysis. To address this gap, we introduce MultiProSE—the first Arabic multilabel dataset comprising 8,000 manually annotated news articles, jointly labeled for propaganda detection, sentiment classification, and fine-grained emotion identification. We extend the ArPro benchmark to establish an open-source multilingual multistage evaluation framework. Leveraging GPT-4o-mini and three BERT variants, we develop reproducible multistage baselines. We publicly release annotation guidelines, source code, and baseline results. This work fills a critical void in Arabic fine-grained opinion mining resources and provides foundational infrastructure for non-English propaganda detection and cross-dimensional stance understanding.
本文提出了一种安全风险评估框架(SRF),通过结合威胁建模、安全分析和基于漏洞关键性的定量风险评估方法,来解决AI生成代码的安全风险评估问题。
This study addresses depression and anxiety among university students stemming from career-related stress by proposing a privacy-preserving, culturally sensitive framework for early mental health risk detection. The approach integrates structured behavioral data with facial emotion cues extracted from interview videos, employing an attention-augmented intermediate fusion neural network for modeling. To enable collaborative training across institutions without sharing raw data, the framework leverages federated learning. Model interpretability is enhanced through Integrated Gradients and SHAP analyses, revealing key behavioral markers—such as gaze avoidance and reduced facial expressivity—that align with established psychological theories. Evaluated on a dataset of Pakistani university students, the model achieves 92.08% accuracy and an F1 score of 89.12%, demonstrating its effectiveness in identifying early signs of psychological distress.
This study systematically evaluates the performance of large language models on Arabic numeral reading tasks—including both Eastern and Western Arabic-Indic digits—across six contextual categories: plain numbers, addresses, dates, quantities, prices, and others. Using a comprehensive benchmark comprising 210 tasks and 59,010 test cases, the authors assess 71 models under four prompting strategies: zero-shot, zero-shot chain-of-thought (CoT), few-shot, and few-shot CoT. For the first time, numerical accuracy is disentangled from instruction-following capability, revealing a wide performance range (14.29%–99.05%) and demonstrating that few-shot CoT improves accuracy by 2.8× over zero-shot. Notably, only six models consistently produce structured outputs, highlighting that high accuracy does not guarantee reliable structured responses; thus, post-processing remains essential for practical deployment.
This study presents the first systematic evaluation of gender and cultural representation biases in mainstream text-to-image models (ImageFX, DALL-E 3, Grok) when generating occupational imagery for Saudi contexts. Using neutral occupation prompts, we generated 1,006 images; two Saudi native annotators performed double-blind labeling on gender, attire, setting, activity, and age, with discrepancies resolved by a senior researcher and supplemented by LLM-assisted analysis. Results reveal pervasive male bias—e.g., 96% male representation in DALL-E 3 outputs—as well as over-masculinization of technical and leadership roles and recurrent cultural misrepresentations in attire and environmental cues. We attribute these biases primarily to cultural homogeneity and societal stereotypes embedded in training data, rather than inherent model limitations. The work establishes the first Arabic-context-specific benchmark and methodological framework for assessing occupational representation fairness in generative AI, advancing equitable AI evaluation practices.
Arabic, the fourth most widely used language online, suffers from a longstanding scarcity of multi-task annotated resources for propaganda, sentiment, and emotion analysis. To address this gap, we introduce MultiProSE—the first Arabic multilabel dataset comprising 8,000 manually annotated news articles, jointly labeled for propaganda detection, sentiment classification, and fine-grained emotion identification. We extend the ArPro benchmark to establish an open-source multilingual multistage evaluation framework. Leveraging GPT-4o-mini and three BERT variants, we develop reproducible multistage baselines. We publicly release annotation guidelines, source code, and baseline results. This work fills a critical void in Arabic fine-grained opinion mining resources and provides foundational infrastructure for non-English propaganda detection and cross-dimensional stance understanding.