APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of current automatic post-editing research, which predominantly relies on sentence-level or synthetic data and fails to support document-level requirements such as terminological consistency and error propagation correction. To bridge this gap, the authors construct the first English–Spanish document-level post-editing dataset in the medical domain, derived from authentic documents used in the UK’s NHS virtual wards and real editing behaviors of professional translators in Trados Studio. The dataset preserves full document structure and computer-assisted translation (CAT) tool context. Initial machine translation drafts were generated using four distinct MT systems and subsequently refined through human post-editing under controlled terminology and quality assurance protocols. The resulting open-resource corpus comprises seven coherent documents totaling 42,000 words, establishing a new benchmark for document-level automatic post-editing and human–machine collaborative translation.
📝 Abstract
Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work on Automatic Post-Editing (APE), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer-Assisted Translation (CAT) workflows. This paper presents the APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings. The corpus contains seven document-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post-edited by professional translators. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE-PE, or LangMark, the dataset preserves document order and CAT-tool context, making it suitable for research on terminology normalisation, correction propagation, and human-in-the-loop translation support. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document-level APE and propagation-aware assistive tools.
Problem

Research questions and friction points this paper is trying to address.

Automatic Post-Editing
Document-Level Translation
Terminology Normalisation
Correction Propagation
Computer-Assisted Translation
Innovation

Methods, ideas, or system contributions that make the work stand out.

document-level APE
correction propagation
CAT-tool context
terminology normalisation
human-in-the-loop translation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marie Escribe
Universitat Politècnica de València
Tharindu Ranasinghe
Tharindu Ranasinghe
Lancaster University, UK
Natural Language ProcessingDeep LearningBenchmarking
A
Amal Haddad Haddad
Universidad de Granada
H
Hansi Hettiarachchi
Lancaster University
Damith Premasiri
Damith Premasiri
PhD Student, Lancaster University
Deep LearningNLPCyber Security