IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过IBBench-Light测试外部指令的任务条件响应,使用配对评估方法和PECA指标,检验了模型在执行和处理指令上的表现。
📝 Abstract
An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, changing the end-of-sequence (EOS) set changes exact paired success from 0/144 to 62/144. A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer. The benchmark measures conditional task and output-contract success. Its task margins and paired count need to be read together with the stopping policy.
Problem

Research questions and friction points this paper is trying to address.

external directives
task-conditioned responses
output contracts
PECA
instruction models
Innovation

Methods, ideas, or system contributions that make the work stand out.

IBBench-Light
Paired Evaluation
Exact-Contract Accuracy (PECA)
Semantic Bases
Directive-Absent Controls
🔎 Similar Papers
No similar papers found.