AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
为解决现有评估框架仅评估AI代理部分功能的问题,AgentAudit提供了一个全面评估规划、工具选择等全生命周期的框架,并通过行为分类和故障归因精确定位故障源。
为解决现有评估框架仅评估AI代理部分功能的问题,AgentAudit提供了一个全面评估规划、工具选择等全生命周期的框架,并通过行为分类和故障归因精确定位故障源。
研究设计了一种声控肌腱驱动的仿生手,通过伺服电机和微控制器实现手指协调运动,解决了手部功能障碍问题。
This study addresses the limitations of current mainstream large language model (LLM) evaluation benchmarks, which are predominantly English- and Western-centric and thus inadequately assess model safety, fairness, and accuracy in India’s multilingual and multicultural context. Building upon the UK AISI’s Inspect AI platform, the authors introduce the first open-source evaluation framework tailored to India’s 22 official languages. The framework encompasses six dimensions: multilingual MMLU, localized bias testing (BharatBBQ), multi-turn jailbreak resistance, cultural knowledge assessment, and safety related to digital public infrastructure (DPI), alongside an LLM-as-judge automated scoring mechanism. Evaluations of five open-source models (8B–32B parameters) reveal that Sarvam-M 24B and Gemma 2 27B both achieve 80% on an Indian fairness index, with Sarvam-M excelling in cultural knowledge and DPI compliance. While all models uniformly reject harmful multilingual prompts (100% refusal rate), their DPI safety scores vary widely (20%–100%).
This study addresses the challenges of dynamically evolving financial fraud and severe class imbalance in digital payment systems by conducting hypothesis-driven exploratory data analysis and feature engineering on the PaySim synthetic dataset, following the CRISP-DM methodology. To mitigate class imbalance, SMOTE oversampling is employed, and hyperparameter optimization is performed via GridSearchCV across multiple classifiers, including logistic regression, decision trees, random forests, and XGBoost. The resulting fraud detection framework achieves significantly enhanced detection performance while maintaining high scalability and robustness, thereby offering FinTech systems an efficient and reliable solution for real-time fraud prevention.
This study addresses the challenges posed by Hinglish—a prevalent Romanized Hindi–English code-mixed variety on Indian social media—whose spelling variations, slang, and out-of-vocabulary terms significantly degrade the performance of conventional monolingual NLP models in sentiment analysis, thereby undermining brand monitoring efforts. To tackle this, the work proposes a high-performance sentiment classification framework specifically designed for Hinglish tweets, which uniquely integrates fine-tuned multilingual BERT (mBERT) with subword tokenization to effectively handle the low-resource nature of code-mixed language. Evaluated on public benchmarks, the approach achieves state-of-the-art accuracy, offering both a deployable, production-ready tool for brand sentiment tracking and a new strong baseline for code-mixing NLP tasks.
为解决现有评估框架仅评估AI代理部分功能的问题,AgentAudit提供了一个全面评估规划、工具选择等全生命周期的框架,并通过行为分类和故障归因精确定位故障源。
研究设计了一种声控肌腱驱动的仿生手,通过伺服电机和微控制器实现手指协调运动,解决了手部功能障碍问题。
This study addresses the limitations of current mainstream large language model (LLM) evaluation benchmarks, which are predominantly English- and Western-centric and thus inadequately assess model safety, fairness, and accuracy in India’s multilingual and multicultural context. Building upon the UK AISI’s Inspect AI platform, the authors introduce the first open-source evaluation framework tailored to India’s 22 official languages. The framework encompasses six dimensions: multilingual MMLU, localized bias testing (BharatBBQ), multi-turn jailbreak resistance, cultural knowledge assessment, and safety related to digital public infrastructure (DPI), alongside an LLM-as-judge automated scoring mechanism. Evaluations of five open-source models (8B–32B parameters) reveal that Sarvam-M 24B and Gemma 2 27B both achieve 80% on an Indian fairness index, with Sarvam-M excelling in cultural knowledge and DPI compliance. While all models uniformly reject harmful multilingual prompts (100% refusal rate), their DPI safety scores vary widely (20%–100%).
This study addresses the challenges of dynamically evolving financial fraud and severe class imbalance in digital payment systems by conducting hypothesis-driven exploratory data analysis and feature engineering on the PaySim synthetic dataset, following the CRISP-DM methodology. To mitigate class imbalance, SMOTE oversampling is employed, and hyperparameter optimization is performed via GridSearchCV across multiple classifiers, including logistic regression, decision trees, random forests, and XGBoost. The resulting fraud detection framework achieves significantly enhanced detection performance while maintaining high scalability and robustness, thereby offering FinTech systems an efficient and reliable solution for real-time fraud prevention.
This study addresses the challenges posed by Hinglish—a prevalent Romanized Hindi–English code-mixed variety on Indian social media—whose spelling variations, slang, and out-of-vocabulary terms significantly degrade the performance of conventional monolingual NLP models in sentiment analysis, thereby undermining brand monitoring efforts. To tackle this, the work proposes a high-performance sentiment classification framework specifically designed for Hinglish tweets, which uniquely integrates fine-tuned multilingual BERT (mBERT) with subword tokenization to effectively handle the low-resource nature of code-mixed language. Evaluated on public benchmarks, the approach achieves state-of-the-art accuracy, offering both a deployable, production-ready tool for brand sentiment tracking and a new strong baseline for code-mixing NLP tasks.