{"id":15,"date":"2024-05-05T19:53:49","date_gmt":"2024-05-05T19:53:49","guid":{"rendered":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/?page_id=15"},"modified":"2026-10-02T13:59:49","modified_gmt":"2026-10-02T13:59:49","slug":"ai-pentesting-literature","status":"publish","type":"page","link":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/research-inquiry\/reliability-testing-of-generative-ai\/ai-pentesting-literature\/","title":{"rendered":"AI &amp; Pentesting Literature"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">This is a list of research papers that report on testing generative AI for pentesting:<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">A thorough testing of ChatGPT&#8217;s capabilities in pentesting<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Deng, G. et al. (2024), \u201cPentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing,\u201d 33rd USENIX Security Symposium, pp. 847\u2013864. <\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Abilities and achievements with Generative AI for Pentesting activities<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Kirmayr, J., Stappen, L. &amp; Andre, E. (2026). <a href=\"https:\/\/aclanthology.org\/2026.acl-long.1886\/\">CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty<\/a>. In <em>Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)<\/em>, pages 40599\u201340618, San Diego, California, United States. Association for Computational Linguistics.<\/li>\n\n\n\n<li>Fang, R., Bindu, R., Gupta, A., Zhan, Q. &amp; Kang, D. (2024)&nbsp;\u2018LLM Agents can Autonomously Hack Websites\u2019. arXiv:2402.06664&nbsp;[cs.CR]<\/li>\n\n\n\n<li>Goyal, D. and Sitaraman Subramanian and Aditya Peela (2024) Hacking, The Lazy Way: LLM Augmented Pentesting.&nbsp; <a href=\"https:\/\/arxiv.org\/abs\/2409.09493\">https:\/\/arxiv.org\/abs\/2409.09493<\/a><\/li>\n\n\n\n<li>Hilario, Azam, S., Sundaram, J., Imran\u00a0Mohammed, K., &amp; Shanmugam, B. (2024). Generative AI for pentesting: the good, the bad, the ugly.\u00a0<em>International Journal of Information Security<\/em>. <a href=\"https:\/\/doi.org\/10.1007\/s10207-024-00835-x\">https:\/\/doi.org\/10.1007\/s10207-024-00835-x<\/a><\/li>\n\n\n\n<li>Iturbe, E., \u2026.. Oscar Llorente-Vazquez, Angel Rego, Erkuden Rios, Nerea Toledo, Unleashing offensive artificial intelligence: Automated attack technique code generation, Computers &amp; Security, Volume 147, 2024, 104077, ISSN 0167-4048, https:\/\/doi.org\/10.1016\/j.cose.2024.104077.<\/li>\n\n\n\n<li>Meyers, B.S., Almassari, S.F., Keller, B.N. &amp; Meneely, A. (2022) \u2018Examining Penetration Tester Behavior in the Collegiate Penetration Testing Competition\u2019,&nbsp;<em>ACM transactions on software engineering and methodology<\/em>, 31(3), pp. 1\u201325.<\/li>\n\n\n\n<li>Raman, R., Calyam, P. &amp; Achuthan, K. (2024) \u2018ChatGPT or Bard: Who is a better Certified Ethical Hacker?\u2019,&nbsp;<em>Computers &amp; security<\/em>, 140p. 103804.<\/li>\n\n\n\n<li>Saber, V. et al (2023) Automated Penetration Testing, A Systematic Review. 2023 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), IEEE.<\/li>\n\n\n\n<li>Saber, V., El Sayad, D., Bahaa-Eldin, A. M. &amp; Fated, Z. T. (2023) Automated Penetration Testing, A Systematic Review. 2023 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), IEEE.<\/li>\n\n\n\n<li>Temara, S. (2023) Maximizing Penetration Testing Success with Effective Reconnaissance Techniques using ChatGPT. https:\/\/doi.org\/10.48550\/arXiv.2307.06391<\/li>\n\n\n\n<li>Wang, P., &amp; D\u2019Cruze, H. (2024). AI-Assisted Pentesting Using ChatGPT-4. In: Latifi, S. (eds) ITNG 2024: 21st International Conference on Information Technology-New Generations. ITNG 2024. Advances in Intelligent Systems and Computing, vol 1456. Springer, Cham. https:\/\/doi.org\/10.1007\/978-3-031-56599-1_9<\/li>\n\n\n\n<li>Wu, L., Zhong, X., Liu, J., Wang, X. (2024). PTGroup: An Automated Penetration Testing Framework Using LLMs and Multiple Prompt Chains. In: Huang, DS., Chen, W., Guo, J. (eds) Advanced Intelligent Computing Technology and Applications. ICIC 2024. Lecture Notes in Computer Science, vol 14870. Springer, Singapore. <a href=\"https:\/\/doi.org\/10.1007\/978-981-97-5606-3_19\">https:\/\/doi.org\/10.1007\/978-981-97-5606-3_19<\/a><\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">These are papers that discuss in general the automation of pentesting<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Al-Hawawreh, M., Aljuhani, A. &amp; Jararweh, Y. (2023) \u2018Chatgpt for cybersecurity: practical applications, challenges, and future directions\u2019,&nbsp;Cluster computing, 26(6), 3421\u20133436.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Happe, A. &amp; Cito, J. (2025) \u2018Benchmarking Practices in LLM-driven Offensive Security:<br>Testbeds, Metrics, and Experiment Design\u2019, Testbeds, Metrics, and Experiment Design, June 2025.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Bhat S. &amp; Varma, V. (2026). <a href=\"https:\/\/aclanthology.org\/2026.findings-acl.1929\/\">All Prompts Are Created Equal? Evaluating Robustness of LLM Judges Against Non-Adversarial Prompt Variations<\/a>. In <em>Findings of the Association for Computational Linguistics: ACL 2026<\/em>, pages 38730\u201338745, San Diego, California, United States. Association for Computational Linguistics.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>This is a list of research papers that report on testing generative AI for pentesting: A thorough testing of ChatGPT&#8217;s capabilities in pentesting Abilities and achievements with Generative AI for Pentesting activities These are papers that discuss in general the &hellip; <a href=\"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/research-inquiry\/reliability-testing-of-generative-ai\/ai-pentesting-literature\/\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":3043,"featured_media":0,"parent":68,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-15","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/pages\/15","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/users\/3043"}],"replies":[{"embeddable":true,"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/comments?post=15"}],"version-history":[{"count":9,"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/pages\/15\/revisions"}],"predecessor-version":[{"id":75,"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/pages\/15\/revisions\/75"}],"up":[{"embeddable":true,"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/pages\/68"}],"wp:attachment":[{"href":"https:\/\/www.staff.ncl.ac.uk\/keerthirajendran\/wp-json\/wp\/v2\/media?parent=15"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}