This is a list of research papers that report on testing generative AI for pentesting:
A thorough testing of ChatGPT’s capabilities in pentesting
- Deng, G. et al. (2024), “PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing,” 33rd USENIX Security Symposium, pp. 847–864.
Abilities and achievements with Generative AI for Pentesting activities
- Kirmayr, J., Stappen, L. & Andre, E. (2026). CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 40599–40618, San Diego, California, United States. Association for Computational Linguistics.
- Fang, R., Bindu, R., Gupta, A., Zhan, Q. & Kang, D. (2024) ‘LLM Agents can Autonomously Hack Websites’. arXiv:2402.06664 [cs.CR]
- Goyal, D. and Sitaraman Subramanian and Aditya Peela (2024) Hacking, The Lazy Way: LLM Augmented Pentesting. https://arxiv.org/abs/2409.09493
- Hilario, Azam, S., Sundaram, J., Imran Mohammed, K., & Shanmugam, B. (2024). Generative AI for pentesting: the good, the bad, the ugly. International Journal of Information Security. https://doi.org/10.1007/s10207-024-00835-x
- Iturbe, E., ….. Oscar Llorente-Vazquez, Angel Rego, Erkuden Rios, Nerea Toledo, Unleashing offensive artificial intelligence: Automated attack technique code generation, Computers & Security, Volume 147, 2024, 104077, ISSN 0167-4048, https://doi.org/10.1016/j.cose.2024.104077.
- Meyers, B.S., Almassari, S.F., Keller, B.N. & Meneely, A. (2022) ‘Examining Penetration Tester Behavior in the Collegiate Penetration Testing Competition’, ACM transactions on software engineering and methodology, 31(3), pp. 1–25.
- Raman, R., Calyam, P. & Achuthan, K. (2024) ‘ChatGPT or Bard: Who is a better Certified Ethical Hacker?’, Computers & security, 140p. 103804.
- Saber, V. et al (2023) Automated Penetration Testing, A Systematic Review. 2023 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), IEEE.
- Saber, V., El Sayad, D., Bahaa-Eldin, A. M. & Fated, Z. T. (2023) Automated Penetration Testing, A Systematic Review. 2023 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), IEEE.
- Temara, S. (2023) Maximizing Penetration Testing Success with Effective Reconnaissance Techniques using ChatGPT. https://doi.org/10.48550/arXiv.2307.06391
- Wang, P., & D’Cruze, H. (2024). AI-Assisted Pentesting Using ChatGPT-4. In: Latifi, S. (eds) ITNG 2024: 21st International Conference on Information Technology-New Generations. ITNG 2024. Advances in Intelligent Systems and Computing, vol 1456. Springer, Cham. https://doi.org/10.1007/978-3-031-56599-1_9
- Wu, L., Zhong, X., Liu, J., Wang, X. (2024). PTGroup: An Automated Penetration Testing Framework Using LLMs and Multiple Prompt Chains. In: Huang, DS., Chen, W., Guo, J. (eds) Advanced Intelligent Computing Technology and Applications. ICIC 2024. Lecture Notes in Computer Science, vol 14870. Springer, Singapore. https://doi.org/10.1007/978-981-97-5606-3_19
These are papers that discuss in general the automation of pentesting
Al-Hawawreh, M., Aljuhani, A. & Jararweh, Y. (2023) ‘Chatgpt for cybersecurity: practical applications, challenges, and future directions’, Cluster computing, 26(6), 3421–3436.
Happe, A. & Cito, J. (2025) ‘Benchmarking Practices in LLM-driven Offensive Security:
Testbeds, Metrics, and Experiment Design’, Testbeds, Metrics, and Experiment Design, June 2025.
Bhat S. & Varma, V. (2026). All Prompts Are Created Equal? Evaluating Robustness of LLM Judges Against Non-Adversarial Prompt Variations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 38730–38745, San Diego, California, United States. Association for Computational Linguistics.