Institution profile

Huaihua University

Academic institutionasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

LATA: A Tool for LLM-Assisted Translation Annotation

Feb 11, 2026

This work addresses the challenge of accurately capturing deep semantic transformations in language pairs with substantial structural divergence, such as Arabic–English, where existing automated tools often fall short. The authors propose an interactive translation annotation method grounded in large language models (LLMs), which employs a templated prompt manager to generate controlled, JSON-formatted sentence segmentation and alignment outputs. By integrating a human-in-the-loop verification process, the approach innovatively combines the scalability of LLMs with the precision of expert annotation. A marginalia-style architecture enables fine-grained labeling of translation strategies, significantly enhancing the quality of parallel corpora for complex language pairs while maintaining annotation efficiency. This framework effectively bridges the gap between fully automated systems and high-quality manual annotation.

0 citationsRead paper

AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts

Dec 25, 2025

Arabic–English parallel corpora are scarce, especially high-quality benchmarks covering challenging domains such as law and literature and supporting many-to-many sentence alignment. To address this gap, we propose the first generative sentence alignment framework tailored to low-resource language pairs: it leverages domain-adapted large language models (LLMs), introduces a novel benchmark—ArabAlign—with a “Hard” subset featuring complex nested structures, ellipsis, and cross-sentence coreference, and defines a fine-grained evaluation protocol. Our method achieves an F1 score of 85.5% on the hard test set, outperforming the prior state of the art by 9 percentage points. Crucially, it is the first approach to robustly model non-one-to-one alignments (e.g., one-to-many, many-to-one) in legal and literary texts. We publicly release a high-quality dataset, open-source implementation, and standardized evaluation tools—filling a critical gap in fine-grained alignment research for low-resource languages.

0 citationsRead paper
Recent publications

Latest Papers

LATA: A Tool for LLM-Assisted Translation Annotation

Feb 11, 2026

This work addresses the challenge of accurately capturing deep semantic transformations in language pairs with substantial structural divergence, such as Arabic–English, where existing automated tools often fall short. The authors propose an interactive translation annotation method grounded in large language models (LLMs), which employs a templated prompt manager to generate controlled, JSON-formatted sentence segmentation and alignment outputs. By integrating a human-in-the-loop verification process, the approach innovatively combines the scalability of LLMs with the precision of expert annotation. A marginalia-style architecture enables fine-grained labeling of translation strategies, significantly enhancing the quality of parallel corpora for complex language pairs while maintaining annotation efficiency. This framework effectively bridges the gap between fully automated systems and high-quality manual annotation.

0 citationsRead paper

AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts

Dec 25, 2025

Arabic–English parallel corpora are scarce, especially high-quality benchmarks covering challenging domains such as law and literature and supporting many-to-many sentence alignment. To address this gap, we propose the first generative sentence alignment framework tailored to low-resource language pairs: it leverages domain-adapted large language models (LLMs), introduces a novel benchmark—ArabAlign—with a “Hard” subset featuring complex nested structures, ellipsis, and cross-sentence coreference, and defines a fine-grained evaluation protocol. Our method achieves an F1 score of 85.5% on the hard test set, outperforming the prior state of the art by 9 percentage points. Crucially, it is the first approach to robustly model non-one-to-one alignments (e.g., one-to-many, many-to-one) in legal and literary texts. We publicly release a high-quality dataset, open-source implementation, and standardized evaluation tools—filling a critical gap in fine-grained alignment research for low-resource languages.

0 citationsRead paper