CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video quality assessment metrics, such as LPIPS and DISTS, overly rely on texture similarity and struggle to capture content-level distortions—such as incorrect objects or blurred text—leading to significant discrepancies with human perception. To address this, this work proposes CodecArena, the first vision-language framework for evaluating video codecs, which formulates quality assessment as conditional contrastive reasoning between reference and reconstructed videos. It introduces a fine-grained evaluation across five dimensions: identity, object, text, texture, and temporal consistency. A novel Facet-GRPO vision-based reinforcement learning algorithm is employed, leveraging auto-generated dimensional cues as weak supervision to enable interpretable and balanced multi-dimensional judgment without manual annotations. Experiments demonstrate that CodecArena achieves state-of-the-art alignment with human subjective ratings across diverse domains, bitrates, and codec types.
📝 Abstract
Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.
Problem

Research questions and friction points this paper is trying to address.

video coding
quality assessment
content fidelity
low bitrate
human perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual reinforcement learning
vision-language framework
fidelity facets
codec quality assessment
Facet-GRPO
J
Jiaye Fu
School of Electronic and Computer Engineering, Peking University
Weiqi Li
Weiqi Li
Ph.D. Candidate, Peking University
Computer VisionSuper ResolutionCompressed Sensing
Q
Qiankun Gao
School of Electronic and Computer Engineering, Peking University
Y
Yanchen Zhao
School of Computer Science, Peking University
Xiandong Meng
Xiandong Meng
University of California Davis
Natural Language Processing LLM Deep Learning
J
Jian Zhang
School of Electronic and Computer Engineering, Peking University
Siwei Ma
Siwei Ma
Peking University
Video Coding and Processing
J
Jiaqi Zhang
School of Computer Science, Peking University