🤖 AI Summary
This study addresses the limitations of current automatic post-editing research, which predominantly relies on sentence-level or synthetic data and fails to support document-level requirements such as terminological consistency and error propagation correction. To bridge this gap, the authors construct the first English–Spanish document-level post-editing dataset in the medical domain, derived from authentic documents used in the UK’s NHS virtual wards and real editing behaviors of professional translators in Trados Studio. The dataset preserves full document structure and computer-assisted translation (CAT) tool context. Initial machine translation drafts were generated using four distinct MT systems and subsequently refined through human post-editing under controlled terminology and quality assurance protocols. The resulting open-resource corpus comprises seven coherent documents totaling 42,000 words, establishing a new benchmark for document-level automatic post-editing and human–machine collaborative translation.
📝 Abstract
Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work on Automatic Post-Editing (APE), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer-Assisted Translation (CAT) workflows. This paper presents the APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings. The corpus contains seven document-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post-edited by professional translators. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE-PE, or LangMark, the dataset preserves document order and CAT-tool context, making it suitable for research on terminology normalisation, correction propagation, and human-in-the-loop translation support. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document-level APE and propagation-aware assistive tools.