Reliability Testing of Generative AI

Generative AI is increasingly being considered for automating complex tasks. However, demonstrating that a model can perform a task is not the same as demonstrating that it can perform it reliably. My MSc research investigated how the reliability of Generative AI can be tested, using penetration testing as the application area. The work developed a framework for designing and conducting reliability tests and for comparing results between tests.

Research problem

There are now many studies demonstrating the ability of Generative AI to perform penetration testing tasks. Much less work has been done on whether these tasks can be performed reliably and there are no established methodologies on testing the reliability of Generative AI systems. Where reliability testing has been reported, various methods have been used and there are different approaches to reporting results. This makes comparison difficult. Testing GenAI poses challenges as systems can produce different outputs for the same input, its internal processes are not transparent, and its performance may depend on prompts, context, model parameters and information encountered during training.

Collated: AI & Pentesting Literature

Findings from MSc project work

We have taken an abductive reasoning approach and used a series of iterative and experimental test designs to develop effective tests. ChatGPT models were used in basic pentesting tasks and where possible, existing work on pentesting using Generative AI models were used as a guide.

A framework was developed to support the design and comparison of reliability tests. It includes guidelines and considerations for test design, a notation for recording and comparing performance, and an approach for dividing complex penetration testing activities into smaller task groups.

The intention is not to provide a single test of whether a GenAI model is reliable. Rather, the framework provides a way of designing tests which can reveal where reliability changes and allow results from different tests to be compared. More information on the following framework elements will be provided provided below:

  • guidelines and recommendations for designing tests
  • a performance benchmarking tool
  • an example of task block separation (pentesting)

[will be available soon]

Future Research Directions

There is considerable scope to extend this work. The framework needs to be tested with other models, larger numbers of tests and more complex penetration testing activities. Further investigation is also needed into the effect of model parameters and other variables on reliability.

A wider problem arising from this work is that existing principles and methods of software testing do not provide sufficient basis for testing Generative AI. Some assumptions made when testing conventional software may not hold for systems whose outputs are variable and whose internal reasoning cannot readily be observed. This work conceptualised and collated a set of testing principles. These need further developing and evaluating with applicability beyond pentesting.

Research suggesting the use of multiagent systems [32], agent-like components [5] and AutoGPT [12] may have the potential to go further beyond scripting interactions mentioned above. The area of multiagent AI is not yet developed enough but may offer a solution to enhance automation with AI.