Too long; didn't solve

📅 2026-04-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how the structural lengths of problem prompts and solutions influence the performance of large language models on mathematical reasoning tasks. By constructing an expert-curated dataset of adversarial math problems and employing correlation analysis alongside a difficulty-normalization methodology, the work reveals—for the first time—a robust association between structural length and the actual difficulty experienced by models. Both prompt length and solution length exhibit significant positive correlations with model failure rates, and even after controlling for intrinsic problem difficulty, they maintain a weak negative correlation with success, indicating that structural length is a key performance factor independent of content complexity. These findings offer a novel perspective on model reasoning bottlenecks and introduce a length-based framework for normalized difficulty analysis.

Technology Category

Application Category

📝 Abstract
Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language models, yet little is known about how their structural properties influence model behaviour. In this work, we investigate two structural length variables, prompt length and solution length, and analyse how they relate to model performance on a newly constructed adversarial dataset of expert-authored mathematics problems. We find that both prompt and solution lengths correlate positively with increased model failure across models. We also include a secondary, exploratory analysis of cross-model disagreement. Under a difficulty-adjusted normalised analysis, both variables retain weak negative associations with realised model separation, slightly stronger for prompt length. Overall, our main robust finding is that structural length is linked to empirical difficulty in this dataset.
Problem

Research questions and friction points this paper is trying to address.

mathematical benchmarks
structural length
prompt length
solution length
model performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

structural length
mathematical reasoning
large language models
adversarial dataset
model failure
L
Lucía M. Cabrera
Instituto Balseiro
I
Isaac Saxton-Knight
Poindexter Labs