Prompt Structure Redistributes, Not Reduces: An Empirical Analysis of Security-Weaknesses in LLM-Generated Python Code

๐Ÿ“… 2026-08-25
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡ๅˆ†ๆž424ไธชPythonไปปๅŠกๅœจไธๅŒๆ็คบ็ป“ๆž„ไธ‹็š„ไปฃ็ ็”Ÿๆˆ๏ผŒๅ‘็Žฐ็ป“ๆž„ๅŒ–ๆ็คบๆ้ซ˜ไบ†ๅˆ่ง„ๆ€งไฝ†ๆœชไธ€่‡ดๅ‡ๅฐ‘ๅฎ‰ๅ…จๅผฑ็‚น๏ผŒไป…้‡ๆ–ฐๅˆ†้…ไบ†้ฃŽ้™ฉใ€‚
๐Ÿ“ Abstract
Large Language Models (LLMs) increasingly generate code from natural-language prompts, making prompt engineering a key mechanism for shaping the security of generated software. Structured and security-oriented prompts are widely used to encourage safer code, yet their effects extend beyond whether detected weaknesses are simply present or absent. Using 424 security-sensitive Python tasks, we generate solutions with GPT-4o and LLaMA 3.1-8B under five prompt variants that progressively add structural and security guidance, and evaluate them with Bandit and CodeQL along two axes: generation compliance and security weakness prevalence, severity, and CWE distributions. Structured prompting substantially reduces refusals (e.g., GPT-4o invalid outputs drop from 338 of 424 to 37-52), enabling large-scale analysis, but security-oriented refinements do not consistently reduce overall weakness prevalence. For GPT-4o, stronger prompts primarily redistribute risk: high-severity findings fall (20.8% to 13.6%) while low-severity findings rise (32% to 43.5%); LLaMA shows weaker, less consistent shifts. We also observe security-driven semantic drift, where stricter prompts silently remove or rewrite explicitly requested unsafe constructs. Overall, prompt structure improves compliance but is an unreliable substitute for robust security controls in LLM-assisted development.
Problem

Research questions and friction points this paper is trying to address.

Security-Weaknesses
LLM-Generated Code
Prompt Engineering
Python Code
Security Controls
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Structure
Security-Weaknesses
Redistribution of Risk
Semantic Drift
Compliance
๐Ÿ’ผ Related Jobs
No related jobs found.