Predicting Program Exit Code with LLMs and Programming Language Semantics

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过程序可执行性预测任务探讨了大型语言模型在理解编程语言语义方面的局限性,发现模型依赖预训练先验而非给定规则。
📝 Abstract
Large language models (LLMs) have shown proficiency in various software engineering tasks, such as code generation and translation. However, a key limitation in their performance may be their (lack of) understanding of programming-language semantics. Even when explicit semantics are given, it remains unclear whether LLMs apply those rules or lean on priors learned during pre-training instead. We study if LLMs lean on priors or given semantics with a novel task--Program Executability Prediction (PrEx)--that asks models to predict whether a program is semantically valid or invalid (and, if invalid, which formal rule it violates) given the program's syntax and operational semantics. Because PrEx requires both valid and invalid programs, we build a dataset with systematically generated invalid transformations derived from valid programs. We evaluate open-source coding LLMs under two semantic formalisms and two semantic shifts across Human-Written, LLM-Translated, and Fuzzer-Generated program splits. Our findings show that LLMs lean on pre-training priors rather than systematically applying the given rules, performing especially poorly on modified semantics and degrading further as program complexity increases. PrEx is available at https://github.com/EngineeringSoftware/prex.
Problem

Research questions and friction points this paper is trying to address.

large language models
programming-language semantics
program executability prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Program Executability Prediction
Programming Language Semantics
Large Language Models
🔎 Similar Papers