CantoneseLLM v2: Reasoning in a Low-Resource Language

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决粤语书面数据稀缺问题,通过CPT、chat-vector合并等方法训练CantoneseLLM v2模型,提高粤语推理能力。
📝 Abstract
Cantonese is widely spoken but remains low-resource in written data, with no large corpus of native Cantonese reasoning traces available for model training. We develop and release CantoneseLLM v2, comprising models based on Qwen3 8B and 30B-A3B. The models are trained through CPT on 784 million Cantonese and Hong Kong-related tokens, chat-vector merging, SFT, DPO, and RLVR. Evaluation across the training stages shows that chat-vector merging transfers instruction following but preserves the donor model's reasoning language, while SFT with limited Cantonese reasoning data substantially shortens or removes reasoning traces and reduces benchmark performance. DPO restores the reasoning-block format, particularly for the 8B model, but recovers only part of the lost performance. The RLVR training with Cantonese language and Traditional Chinese scripts as multiplicative constraints introduced Cantonese language alignment and restored the lost performance. The 30B-A3B model reaches 73.16 on HKCanto-Eval, within 1.20 points of its merged checkpoint, while retaining the Cantonese reasoning behaviour absent from that checkpoint. We release the model checkpoints, the training environments, and a thirteen-year Traditional Chinese Common Crawl dataset. The models can be accessed at https://huggingface.co/collections/hon9kon9ize/cantonesellm-v20
Problem

Research questions and friction points this paper is trying to address.

Cantonese
Low-Resource Language
Reasoning Traces
Innovation

Methods, ideas, or system contributions that make the work stand out.

CantoneseLLM v2
chat-vector merging
SFT
DPO
RLVR
💼 Related Jobs
No related jobs found.
T
Tsz Chung Cheng
Kyushu University
C
Chung Shing Cheng
hon9kon9ize
C
Chaak Ming Lau
The Education University of Hong Kong
C
Cheuk Hei Chong
V otee AI