Check The Scoreboard: An Analysis of Scoring Schemes on Multiple-Choice Evaluation

๐Ÿ“… 2026-08-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡้‡‡็”จๅ…ญ็งๆ•™่‚ฒๅฏๅ‘็š„่ฏ„ๅˆ†ๆ–นๆกˆ๏ผŒๅฆ‚ๅนฒๆ‰ฐ้กนๆŽ’้™คใ€ๅผƒๆƒ็ญ‰๏ผŒๆฅ่ฏ„ไผฐๅคš้€‰้ข˜ๅ›ž็ญ”ไธญ้™คๅ‡†็กฎ็އๅค–็š„่ƒฝๅŠ›๏ผŒๆญ็คบไธๅŒๆจกๅž‹็š„็‹ฌ็‰น่ƒฝๅŠ›ใ€‚
๐Ÿ“ Abstract
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.
Problem

Research questions and friction points this paper is trying to address.

Multiple-Choice Question Answering
Scoring Schemes
Accuracy
Educational Testing
Model Capabilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Alternative Scoring Schemes
Distractor Elimination
Abstention
Confidence Calibration
Self-Correction
๐Ÿ”Ž Similar Papers
No similar papers found.