Rubric-to-Code Credit Assignment for Reinforcement Learning

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Rubric-to-Code Credit Assignment (RCCA)方法,通过将功能性反馈转化为局部优化信号来改进交互式Web应用的代码生成,显著提高了模型在多个基准测试上的表现。
📝 Abstract
Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requirements, each often tied to localized code regions such as event handlers, state updates, DOM fragments, or CSS selectors. Standard GRPO collapses these structured outcomes into a single sequence-level reward and applies the resulting advantage uniformly to all tokens, weakening credit assignment. We propose \textbf{Rubric-to-Code Credit Assignment} (RCCA), a reinforcement learning framework that converts rubric-level functional feedback into localized optimization signals over generated code. RCCA builds training tasks around explicit functional rubrics, uses a hierarchical reward to separate format, source-code, runtime, and functional failures, and aligns evaluator-generated textual attributions with responsible code spans and generated tokens. The resulting model, \textbf{Ling-RCCA-Flash}, scores 41.25 on MiniAppBench, improving Ling-3.0-Flash by 32.20 points and slightly surpassing Claude Opus 4.5. It also reaches 76.19 on ArtifactsBench, improving the SFT model by 4.48 points and establishing a new top score under the official ArtifactsBench leaderboard setting by surpassing the GPT-5 score by 3.64 points, suggesting transferable implementation-level gains.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Interactive Web Application
Code Generation
Functional Requirements
Credit Assignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Credit Assignment
Hierarchical Reward
Functional Feedback
Code Generation
🔎 Similar Papers
Rui Jin
Rui Jin
Cold and Arid Regions Environment and Engineering Research Institute, Chinese Academy of Sciences
surface freeze/thaw cyclessoil moisturemicrowave remote snesingWSNdata assimilation
J
Jikai Chen
Inclusion AI, Ant Group
Y
Yihan Chen
Inclusion AI, Ant Group
H
Hao Zhou
Inclusion AI, Ant Group
D
Demin Zhu
Inclusion AI, Ant Group
K
Kaichen Yang
Inclusion AI, Ant Group
D
Dong Wang
Inclusion AI, Ant Group
Chenyi Zhuang
Chenyi Zhuang
AIST, AIRC
machine learning