What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了从ASR转录中检查古兰经背诵错误的问题,通过人工标注100个录音案例,并使用可执行评估器评分,以区分不同类型的错误。
📝 Abstract
Checking Quran recitation from an ASR transcript requires distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences. We report a completed human annotation of 100 production recording cases: 348 scored units and 162 localized events across ten combined labels. An executable evaluator scores labels and word positions together. A plain diff reaches label-aware F1 0.525 and localization F1 0.826; adapted production cleaner/alignment components reach 0.518 and 0.786, with exact-span F1 0.505 for both. Correcting the adapter's word coordinates recovers all five annotated repetition events, showing why annotation interfaces must be checked before interpreting baseline failures. In a preliminary pilot, eight single 20-minute runs across three coding agents and eight models span label-aware F1 0.143 to 0.892: seven land far above every baseline, and one collapses below the naive diff from a missing normalization step. Across the six, 970 of 972 gold-event instances draw an overlapping prediction, so what remains is not detection but convention: span extent, and the labels whose boundary is stipulated by adjudication rather than visible in the text. Seven of 162 events defeat all six same-day runs, five of them one orthographic rule, and the strongest run still misses the same ones. No run annotated before building, so the pilot measures the algorithm half of the task only.
Problem

Research questions and friction points this paper is trying to address.

Quran recitation
ASR transcript
human annotation
mistake identification
event types
Innovation

Methods, ideas, or system contributions that make the work stand out.

executable evaluator
Quran recitation
ASR transcript
annotation interface
baseline failures
💼 Related Jobs
No related jobs found.