Institution profile

2077AI

Industry research
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values

Apr 07, 2025

Existing Chinese preference datasets suffer from limited scale, narrow domain coverage, weak validation, and bottlenecks in manual annotation. To address these issues, this work introduces the first large-scale, high-quality Chinese preference dataset—comprising 1.009 million preference pairs—spanning six domains including dialogue, coding, and mathematics, specifically designed to support value alignment of large language models (LLMs). We propose a fully automated LLM-based preference annotation pipeline, eliminating human intervention. We release the first open-source Chinese Reward Model (CRM) and Chinese Reward Benchmark (CRBench); CRM achieves GPT-4o-level performance in low-quality sample detection. Trained on responses from 15 mainstream LLMs using multi-stage crawling and filtering, CRM has 8B parameters. We establish a dual-evaluation framework—AlignBench and CRBench—and demonstrate 2–12% alignment improvement on Qwen2/2.5 and Infinity-Instruct series. All code and data are publicly available.

0 citationsRead paper
Recent publications

Latest Papers

COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values

Apr 07, 2025

Existing Chinese preference datasets suffer from limited scale, narrow domain coverage, weak validation, and bottlenecks in manual annotation. To address these issues, this work introduces the first large-scale, high-quality Chinese preference dataset—comprising 1.009 million preference pairs—spanning six domains including dialogue, coding, and mathematics, specifically designed to support value alignment of large language models (LLMs). We propose a fully automated LLM-based preference annotation pipeline, eliminating human intervention. We release the first open-source Chinese Reward Model (CRM) and Chinese Reward Benchmark (CRBench); CRM achieves GPT-4o-level performance in low-quality sample detection. Trained on responses from 15 mainstream LLMs using multi-stage crawling and filtering, CRM has 8B parameters. We establish a dual-evaluation framework—AlignBench and CRBench—and demonstrate 2–12% alignment improvement on Qwen2/2.5 and Infinity-Instruct series. All code and data are publicly available.

0 citationsRead paper