Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
研究评估了尼日利亚761名医疗专业人员对AI的准备情况,发现尽管意识高但知识和准备不足,提出需加强培训、基础设施投资等措施。
研究评估了尼日利亚761名医疗专业人员对AI的准备情况,发现尽管意识高但知识和准备不足,提出需加强培训、基础设施投资等措施。
This study addresses the limited real-world applicability of existing agricultural vision models, which are predominantly trained on idealized datasets and struggle in the complex, dynamic field conditions prevalent in under-resourced regions such as Africa. To bridge this gap, the authors present a systematic evaluation of six state-of-the-art object detection models—YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR—on AgriAISeg, the first multi-crop, multi-challenge dataset collected directly from African farmlands. Experimental results demonstrate that RT-DETR achieves the highest performance with an mAP@0.5:0.95 of 0.624, while YOLO-family models exhibit consistently strong accuracy and training efficiency. In contrast, Faster R-CNN suffers significant performance degradation in complex scenarios. This work provides the first empirical evidence of the varying suitability of modern detectors in authentic African agricultural settings, offering critical guidance for deploying AI solutions in resource-constrained environments.
Current safety alignment of large language models is predominantly based on English, and their cross-lingual generalization to low-resource languages remains poorly understood, posing potential risks. This work introduces the LoDNA dataset, comprising both literal translations and culturally localized prompts, to systematically evaluate the transferability of safety mechanisms across four African languages. We propose a probing method grounded in the geometric structure of the model’s latent space to analyze internal representations underlying refusal behaviors. Our study reveals, for the first time, significant limitations in cross-lingual safety alignment: in most language–model combinations, harmful prompts retain less than 10% of the refusal signal observed in English, indicating that semantic alignment does not ensure consistent safety routing. These findings challenge the assumption of a language-invariant harm manifold.
This study addresses the challenges of high cost, limited spatial coverage, and susceptibility to adverse weather that plague conventional in situ coastal wave observation methods, which hinder efficient acquisition of wave parameters. To overcome these limitations, the authors propose a deep learning framework that leverages monocular coastal video to jointly estimate five key wave parameters—significant wave height, maximum wave height, peak period, zero-crossing period, and wave direction—under data-scarce conditions. The approach integrates a self-supervised V-JEPA vision transformer, a SlowFast dual-stream temporal encoder, and Farneback optical flow, while incorporating Airy dispersion relation constraints to enforce physical consistency. Validated with only six annotated scenes and accelerated via high-performance computing, the model achieves Pearson correlation coefficients ranging from 0.451 to 0.832 across all five parameters, demonstrating both feasibility and strong cross-site generalization capability.
This study addresses the limitations of traditional license plate recognition systems, which rely on multi-stage pipelines combining YOLO and OCR and suffer from performance degradation, high computational costs, and dependence on annotated data in unstructured environments. For the first time, this work introduces vision-language models (VLMs) to license plate recognition in Nigeria’s challenging real-world conditions, proposing a zero-shot, end-to-end unified framework that enables direct recognition without task-specific training. The authors evaluate leading VLMs—including Gemini 2.0 Flash, Qwen2.5-VL, GPT-4o, Claude 4 Sonnet, and Llama 3.2 Vision—on a dataset of 88 real-world images. Results demonstrate that Gemini and Qwen significantly outperform other models, achieving lower character error rates and confirming the feasibility and robustness of VLMs as a viable alternative to conventional YOLO+OCR pipelines.
研究评估了尼日利亚761名医疗专业人员对AI的准备情况,发现尽管意识高但知识和准备不足,提出需加强培训、基础设施投资等措施。
This study addresses the limited real-world applicability of existing agricultural vision models, which are predominantly trained on idealized datasets and struggle in the complex, dynamic field conditions prevalent in under-resourced regions such as Africa. To bridge this gap, the authors present a systematic evaluation of six state-of-the-art object detection models—YOLOv5, YOLOv8, YOLO11, YOLO26, Faster R-CNN, and RT-DETR—on AgriAISeg, the first multi-crop, multi-challenge dataset collected directly from African farmlands. Experimental results demonstrate that RT-DETR achieves the highest performance with an mAP@0.5:0.95 of 0.624, while YOLO-family models exhibit consistently strong accuracy and training efficiency. In contrast, Faster R-CNN suffers significant performance degradation in complex scenarios. This work provides the first empirical evidence of the varying suitability of modern detectors in authentic African agricultural settings, offering critical guidance for deploying AI solutions in resource-constrained environments.
Current safety alignment of large language models is predominantly based on English, and their cross-lingual generalization to low-resource languages remains poorly understood, posing potential risks. This work introduces the LoDNA dataset, comprising both literal translations and culturally localized prompts, to systematically evaluate the transferability of safety mechanisms across four African languages. We propose a probing method grounded in the geometric structure of the model’s latent space to analyze internal representations underlying refusal behaviors. Our study reveals, for the first time, significant limitations in cross-lingual safety alignment: in most language–model combinations, harmful prompts retain less than 10% of the refusal signal observed in English, indicating that semantic alignment does not ensure consistent safety routing. These findings challenge the assumption of a language-invariant harm manifold.
This study addresses the challenges of high cost, limited spatial coverage, and susceptibility to adverse weather that plague conventional in situ coastal wave observation methods, which hinder efficient acquisition of wave parameters. To overcome these limitations, the authors propose a deep learning framework that leverages monocular coastal video to jointly estimate five key wave parameters—significant wave height, maximum wave height, peak period, zero-crossing period, and wave direction—under data-scarce conditions. The approach integrates a self-supervised V-JEPA vision transformer, a SlowFast dual-stream temporal encoder, and Farneback optical flow, while incorporating Airy dispersion relation constraints to enforce physical consistency. Validated with only six annotated scenes and accelerated via high-performance computing, the model achieves Pearson correlation coefficients ranging from 0.451 to 0.832 across all five parameters, demonstrating both feasibility and strong cross-site generalization capability.
This study addresses the limitations of traditional license plate recognition systems, which rely on multi-stage pipelines combining YOLO and OCR and suffer from performance degradation, high computational costs, and dependence on annotated data in unstructured environments. For the first time, this work introduces vision-language models (VLMs) to license plate recognition in Nigeria’s challenging real-world conditions, proposing a zero-shot, end-to-end unified framework that enables direct recognition without task-specific training. The authors evaluate leading VLMs—including Gemini 2.0 Flash, Qwen2.5-VL, GPT-4o, Claude 4 Sonnet, and Llama 3.2 Vision—on a dataset of 88 real-world images. Results demonstrate that Gemini and Qwen significantly outperform other models, achieving lower character error rates and confirming the feasibility and robustness of VLMs as a viable alternative to conventional YOLO+OCR pipelines.