🤖 AI Summary
This study investigates the development of students’ critical evaluation capabilities in AI-assisted translation pedagogy by engaging them in comparative assessments of outputs from general-purpose large language models and online machine translation systems. Participants performed post-editing tasks and justified their decisions using both automatic metrics (e.g., BLEU) and human evaluations focusing on fluency and accuracy. Findings reveal that students primarily based their judgments on multidimensional criteria—including terminological precision, linguistic naturalness, and anticipated editing effort—and often preferred translations that diverged from rankings suggested by automatic scores, demonstrating a nuanced, critical judgment that transcends quantitative metrics. This work challenges conventional paradigms of system evaluation by uncovering the complexity and pedagogical significance of learners’ subjective assessment logic in authentic instructional contexts.
📝 Abstract
Drawing on 23 anonymized student pro-jects from a fourth-year Machine Transla-tion and Post-editing course in a BA-level translation programme, this paper exam-ines how structured comparison of gen-eral-purpose LLMs and online MT sys-tems can elicit evaluative judgement in AI-mediated translation. Students translat-ed short specialised English Wikipedia texts into Catalan or Spanish, generated four system outputs, evaluated them using automatic metrics and human adequa-cy/fluency assessment, selected one output for post-editing, and justified their deci-sion in written reports. Descriptive counts are reported for all 23 projects, while qualitative interpretation is based on the 22 cases accompanied by written reports. Results show that students did not treat automatic metrics as final authority: final post-editing selections often diverged from metric rankings and were justified through adequacy, fluency, terminology, naturalness, and expected post-editing ef-fort. The study therefore does not bench-mark systems under controlled conditions; it analyses how students justified system choice within an authentic classroom as-signment.