Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文使用强化学习和大型语言模型提高加泰罗尼亚语文本简化能力,通过引入新的奖励函数并结合SARI指标和特定惩罚项来优化模型。
📝 Abstract
Although automatic text simplification (ATS) is critical for accessibility, its progress has not matched the rapid evolution of broader natural language processing techniques. This paper investigates the application of reinforcement learning (RL) to improve the quality of ATS for low-resource languages using Large Language Models (LLMs). The paper introduces a novel reward function, designed to guide LLMs toward a targeted simplification style with Group Relative Policy Optimization (GRPO), that combines the SARI metric with specific penalty components. The effectiveness of GRPO with this reward function is motivated and demonstrated by post-training IberianLLM-7B-Instruct on the ASSET dataset. After post-training on the English ASSET, the model's ATS performance improves on two curated Catalan benchmarks while also successfully suppressing previously observed negative behaviors. Cross-lingual transfer learning is explored by translating ASSET into Catalan and Spanish and post-training the model on each version, but these fail to show a significant improvement on the out-of-domain benchmark.
Problem

Research questions and friction points this paper is trying to address.

Automatic Text Simplification
Low-resource Languages
Reinforcement Learning
Large Language Models
Catalan
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Group Relative Policy Optimization (GRPO)
SARI metric
Large Language Models (LLMs)
🔎 Similar Papers
No similar papers found.