Abstract
This study examines how nine AI code-generation models use contextual information when translating natural-language descriptions into software exploits. Experiments on real-world shellcodes compare fine-tuned and instruction-tuned models under related, missing, and irrelevant context. The findings show that encoder-decoder models benefit most from relevant context, while decoder-only models can gain indirectly from unrelated context and instruction-tuned LLMs struggle to exploit context consistently. The results motivate task-specific fine-tuning and deliberate context selection for security-sensitive code generation.
Type
Publication
Empirical Software Engineering