LAION-Mobile: Evaluating Deepfake Detectors On One Million Smartphone Photos
研究通过构建包含100万张智能手机照片的数据集LAION-Mobile,评估了12种深度伪造检测器在现代手机摄影上的表现,发现这些检测器对现代AI内容的识别效果不佳。
研究通过构建包含100万张智能手机照片的数据集LAION-Mobile,评估了12种深度伪造检测器在现代手机摄影上的表现,发现这些检测器对现代AI内容的识别效果不佳。
研究使用深度神经网络解决表格内预测问题,通过自监督学习方法和合成数据评估MLP、Resnet和Transformer三种架构性能。
This work addresses two structural failure modes in AI coding agents—context explosion and silent specification-code drift—that undermine development efficiency and code reliability. To tackle these challenges, the paper proposes a lightweight framework that uniquely integrates classical software engineering principles, including information hiding, the C4 model, and Architecture Decision Records (ADRs), into a unified, machine-enforceable architecture. The framework employs a machine-readable specification graph, an ownership-path-based Spine context assembler, a vertical-slice growth protocol, and a drift-gating mechanism to jointly ensure consistency between specifications and generated code. By constraining context expansion and mandating drift detection and correction prior to code merge, the approach significantly enhances the reliability and maintainability of AI-generated code.
This study addresses the inefficiency and error-proneness of manual grading of handwritten answer sheets, particularly in complex scenarios involving answer misalignment, strikeovers, or connected writing, where existing automated methods suffer from insufficient accuracy and compromised fairness. The work proposes the first application of general-purpose vision-language models (VLMs) to handwritten answer recognition, leveraging holistic page-level semantic understanding and prompt engineering guided by reference answers to achieve high-precision single-character judgment. A novel fairness-oriented evaluation framework is introduced, explicitly distinguishing between false negatives and false positives to significantly reduce false negative rates—errors disproportionately disadvantageous to students. Evaluated on 61 anonymized exam papers comprising 3,141 answer fields, the method achieves 98.4% overall accuracy with a false negative rate of only 0.58%; among all papers, only three received lower scores, and these discrepancies are detectable through a student self-review step.
This study addresses the pedagogical narrowing, technical limitations, and compliance challenges associated with fully or partially digitized summative assessments in large-scale higher education examinations. To reconcile the instructional value of open-ended, problem-oriented questions with the need for scalable grading, the authors propose a hybrid e-assessment approach that retains paper-based handwritten responses within a structured answer format. This method leverages a two-stage verification pipeline: first, visual large language models and handwriting recognition technologies automatically extract handwritten content from standardized response tables; second, the extracted answers undergo semantic comparison against reference answer keys. The proposed framework preserves the educational benefits of open-response items while substantially reducing recognition errors, thereby enhancing the accuracy, fairness, and scalability of summative assessment in higher education.
研究通过构建包含100万张智能手机照片的数据集LAION-Mobile,评估了12种深度伪造检测器在现代手机摄影上的表现,发现这些检测器对现代AI内容的识别效果不佳。
研究使用深度神经网络解决表格内预测问题,通过自监督学习方法和合成数据评估MLP、Resnet和Transformer三种架构性能。
This work addresses two structural failure modes in AI coding agents—context explosion and silent specification-code drift—that undermine development efficiency and code reliability. To tackle these challenges, the paper proposes a lightweight framework that uniquely integrates classical software engineering principles, including information hiding, the C4 model, and Architecture Decision Records (ADRs), into a unified, machine-enforceable architecture. The framework employs a machine-readable specification graph, an ownership-path-based Spine context assembler, a vertical-slice growth protocol, and a drift-gating mechanism to jointly ensure consistency between specifications and generated code. By constraining context expansion and mandating drift detection and correction prior to code merge, the approach significantly enhances the reliability and maintainability of AI-generated code.
This study addresses the inefficiency and error-proneness of manual grading of handwritten answer sheets, particularly in complex scenarios involving answer misalignment, strikeovers, or connected writing, where existing automated methods suffer from insufficient accuracy and compromised fairness. The work proposes the first application of general-purpose vision-language models (VLMs) to handwritten answer recognition, leveraging holistic page-level semantic understanding and prompt engineering guided by reference answers to achieve high-precision single-character judgment. A novel fairness-oriented evaluation framework is introduced, explicitly distinguishing between false negatives and false positives to significantly reduce false negative rates—errors disproportionately disadvantageous to students. Evaluated on 61 anonymized exam papers comprising 3,141 answer fields, the method achieves 98.4% overall accuracy with a false negative rate of only 0.58%; among all papers, only three received lower scores, and these discrepancies are detectable through a student self-review step.
This study addresses the pedagogical narrowing, technical limitations, and compliance challenges associated with fully or partially digitized summative assessments in large-scale higher education examinations. To reconcile the instructional value of open-ended, problem-oriented questions with the need for scalable grading, the authors propose a hybrid e-assessment approach that retains paper-based handwritten responses within a structured answer format. This method leverages a two-stage verification pipeline: first, visual large language models and handwriting recognition technologies automatically extract handwritten content from standardized response tables; second, the extracted answers undergo semantic comparison against reference answer keys. The proposed framework preserves the educational benefits of open-response items while substantially reducing recognition errors, thereby enhancing the accuracy, fairness, and scalability of summative assessment in higher education.