SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition
为解决AI生成的教学幻灯片视觉效果与教学效果不匹配的问题,通过构建SLATE基准评估系统,采用预后测设计和学习者知识获取来评价其有效性。
为解决AI生成的教学幻灯片视觉效果与教学效果不匹配的问题,通过构建SLATE基准评估系统,采用预后测设计和学习者知识获取来评价其有效性。
Existing robotic reward models struggle to simultaneously maintain pointwise scoring accuracy and pairwise preference consistency in long-horizon tasks, leading to training noise and performance degradation. This work proposes Preference-Ordered Isotonic Score Editing (POISE), a novel method that achieves conflict-free alignment between these two signal types for the first time, effectively resolving the score-preference reversal problem. Evaluated on a unified four-paradigm dataset and leveraging a vision-language model with video question-answering supervision and the TrustJudge reasoning aggregation strategy, Qwen3-VL-4B calibrated by POISE attains a reward accuracy of 77.96% and improves score-preference consistency to 71.90%. Further integration of TrustJudge elevates the overall performance to 78.57%, surpassing the teacher model.
This work addresses the challenges of decoding high-dimensional Chinese sentences from non-invasive electroencephalography (EEG), including the large character set, substantial inter-subject variability, and low signal-to-noise ratio. To this end, the authors propose EEGAlign, a novel framework that achieves, for the first time, large-vocabulary Chinese sentence decoding from non-invasive EEG. The method innovatively employs contrastive learning to jointly align EEG signals with both textual semantics (via BGE-M3) and acoustic features (via wav2vec 2.0), followed by connectionist temporal classification (CTC) for sequence decoding. Evaluated on the ChineseEEG-2 dataset, EEGAlign attains state-of-the-art performance, achieving 82.37% top-1 accuracy in the speech-reading task and 41.43% in the passive listening task among 101 candidate sentences, thereby demonstrating the complementary benefits of dual-axis alignment in enhancing sentence-level discriminability and temporal resolution.
This study addresses the gap in value alignment evaluation of large language models (LLMs), which has predominantly reflected Western perspectives and overlooked Indonesian indigenous values, particularly Pancasila. The authors introduce the first benchmark dataset grounded in Pancasila’s five core principles—belief in one God, just and civilized humanity, national unity, democracy guided by wisdom, and social justice—comprising 1,834 moral dilemmas derived from Indonesian news sources. These items were curated by native speakers and validated through multi-annotator voting. Evaluating 50 prominent LLMs using Probability Matching Score (PMS) and Majority Vote Agreement Score (MVAS), the study reveals that all models score below 0.5 on PMS and below 0.72 on MVAS, with notably poor performance on religious and unity-related dimensions, underscoring a significant deficiency in their understanding of Indonesian cultural and ethical norms.
Existing defenses struggle against multi-turn jailbreaking attacks due to their reliance on costly retraining, degradation of model utility, or limitation to single-turn analysis, which fails to capture the temporal accumulation of risk across dialogue turns. This work proposes the first training-free defense framework specifically designed for multi-turn jailbreaking scenarios, introducing an explicit mechanism to model temporal risk accumulation. By integrating decay modulation and trend awareness, the framework dynamically fuses safety signals from each turn’s input, historical intent evolution, and model outputs. It comprises a turn-level risk evaluator, a historical context analyzer, a response evaluator, and a dynamic decision module. Evaluated on mainstream large language models, the approach reduces attack success rates to 0.2%–4.0% with no more than 1.5% utility loss, and effectively blocks over 70% of attacks only from the second turn onward.
为解决AI生成的教学幻灯片视觉效果与教学效果不匹配的问题,通过构建SLATE基准评估系统,采用预后测设计和学习者知识获取来评价其有效性。
Existing robotic reward models struggle to simultaneously maintain pointwise scoring accuracy and pairwise preference consistency in long-horizon tasks, leading to training noise and performance degradation. This work proposes Preference-Ordered Isotonic Score Editing (POISE), a novel method that achieves conflict-free alignment between these two signal types for the first time, effectively resolving the score-preference reversal problem. Evaluated on a unified four-paradigm dataset and leveraging a vision-language model with video question-answering supervision and the TrustJudge reasoning aggregation strategy, Qwen3-VL-4B calibrated by POISE attains a reward accuracy of 77.96% and improves score-preference consistency to 71.90%. Further integration of TrustJudge elevates the overall performance to 78.57%, surpassing the teacher model.
This work addresses the challenges of decoding high-dimensional Chinese sentences from non-invasive electroencephalography (EEG), including the large character set, substantial inter-subject variability, and low signal-to-noise ratio. To this end, the authors propose EEGAlign, a novel framework that achieves, for the first time, large-vocabulary Chinese sentence decoding from non-invasive EEG. The method innovatively employs contrastive learning to jointly align EEG signals with both textual semantics (via BGE-M3) and acoustic features (via wav2vec 2.0), followed by connectionist temporal classification (CTC) for sequence decoding. Evaluated on the ChineseEEG-2 dataset, EEGAlign attains state-of-the-art performance, achieving 82.37% top-1 accuracy in the speech-reading task and 41.43% in the passive listening task among 101 candidate sentences, thereby demonstrating the complementary benefits of dual-axis alignment in enhancing sentence-level discriminability and temporal resolution.
This study addresses the gap in value alignment evaluation of large language models (LLMs), which has predominantly reflected Western perspectives and overlooked Indonesian indigenous values, particularly Pancasila. The authors introduce the first benchmark dataset grounded in Pancasila’s five core principles—belief in one God, just and civilized humanity, national unity, democracy guided by wisdom, and social justice—comprising 1,834 moral dilemmas derived from Indonesian news sources. These items were curated by native speakers and validated through multi-annotator voting. Evaluating 50 prominent LLMs using Probability Matching Score (PMS) and Majority Vote Agreement Score (MVAS), the study reveals that all models score below 0.5 on PMS and below 0.72 on MVAS, with notably poor performance on religious and unity-related dimensions, underscoring a significant deficiency in their understanding of Indonesian cultural and ethical norms.
Existing defenses struggle against multi-turn jailbreaking attacks due to their reliance on costly retraining, degradation of model utility, or limitation to single-turn analysis, which fails to capture the temporal accumulation of risk across dialogue turns. This work proposes the first training-free defense framework specifically designed for multi-turn jailbreaking scenarios, introducing an explicit mechanism to model temporal risk accumulation. By integrating decay modulation and trend awareness, the framework dynamically fuses safety signals from each turn’s input, historical intent evolution, and model outputs. It comprises a turn-level risk evaluator, a historical context analyzer, a response evaluator, and a dynamic decision module. Evaluated on mainstream large language models, the approach reduces attack success rates to 0.2%–4.0% with no more than 1.5% utility loss, and effectively blocks over 70% of attacks only from the second turn onward.