PAPER / ARXIV:2609.18298
Shafique, A.; Miller, B.; Heymann, E.
RESUMO
Agentic-AI based software development offers the promise of faster completion of software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question on three practices. First, we applied a typical best-practices agentic AI workflow for software development. Second, our target programs were ten well-known, release-quality human-written Linux utility programs so that we could compare the AI-generated code against a concrete ground truth. Third, we measure reliability on widely used testing technique, fuzz random testing. For this testing, we used both classic black box, generational testing and modern coverage guided (gray box, mutational) testing using AFL++. We found that versions of AI-generated code are typically as often more reliable than the latest human-generated versions of programs. While AI-generated code did have some failures, they were less common than from standard repositories. Interestingly, AI-generated code was likely to have failures such as memory errors (such as buffer overflows) but more likely hangs than infinite loops. In addition, we verified that generating robust agentic AI based software requires careful practice and human supervision. The quality of the software is highly dependent on the prompts and skills used, directing the human process responds. We also demonstrated (with agentic AI its prompts and skills) can become a specification of leads to cost-effective sustainability of software.
NO MESMO MAPA