Generative artificial intelligence can draft, explain, predict and personalise at a scale and speed we have not seen before. Yet beneath its fluent and confident outputs is a characteristic that changes some familiar assumptions about risk and assurance.
Traditionally, when decisions affect sensitive data or vulnerable people, we expect a high degree of certainty from the software we approve and from the vendors who provide it. That expectation is harder to sustain with new AI-enabled educational tools. In reviewing software vendors, we cannot expect certainty where certainty cannot reasonably be provided.
For school leaders navigating this landscape, the challenge when approving AI tools is to ascertain what can be known, where uncertainty remains, and whether the residual risk is acceptable.
The black box problem
At the heart of generative AI are systems built on very large numbers of learned numerical parameters, adjusted during training to recognise patterns and generate responses.
Unlike conventional software built around explicit rules and traceable logic, the internal basis for an individual generative AI output is not fully interpretable or straightforward to trace. Developers can inspect the architecture, training process and system behaviour, but they cannot provide a complete human-readable explanation for every output.
This is one reason these systems are often described as a “black box”.
That has several implications:
- Generative AI is probabilistic. The same or similar prompts can produce different outputs, and those outputs can appear convincing while still being inaccurate, incomplete or inappropriate.
- It is also context sensitive. Behaviour can change with relatively small differences in wording, underlying data, system configuration, model updates or the way a tool is integrated with other systems.
- And its internal representations are not equivalent to human reasoning. Even where developers can explain how the system was designed and tested, that does not mean they can fully explain why each output was produced.
“Assurance means having sufficient evidence to be reasonably confident that the risks are understood and can be appropriately managed.”
Redefining assurance
For leaders, this means assurance cannot mean asking a provider to guarantee that an AI system will always behave correctly.
Without a process designed to deal explicitly with uncertainty, two understandable responses are to avoid the technology because the uncertainty feels unacceptable, or to rely too heavily on assurance from the supplier or IT function. Neither response, on its own, resolves the governance problem.
A realistic standard is to ask vendors for information that can help with risk assessment. Assurance now means having sufficient evidence to be reasonably confident that the risks are understood and can be appropriately managed.
The question: “Is this safe for our school?” can be unpacked to a set of questions:
- What do we know? What evidence do we have?
- Where are the gaps? Where does uncertainty remain?
- What controls can we put in place?
- What residual risk will still remain?
The leadership decision is whether the residual risk is understood well enough to make a deliberate and proportionate decision.
Subsequently, AI approval includes two related stages:
- gaining proportionate assurance, and
- deciding whether the remaining risk is acceptable.
What schools can reasonably expect from suppliers
The Department for Education’s Generative AI: product safety standards set out expectations for developers and suppliers of AI products used in education.
They give school and trust leaders a useful basis for asking questions about areas such as safety, transparency, testing, data protection and the management of foreseeable harms. Suppliers have an important role in testing their products, providing appropriate safeguards and explaining the limitations of their systems.
But they cannot determine whether every use of their product will be risk-free in a school context. That decision remains with the organisation adopting the technology.
Leaders still need to ask questions like:
- Where is the risk greatest, and who may be most vulnerable?
- What human oversight will sit alongside the tool?
- What training, policy or restrictions are needed?
- How will use of the tool be reviewed?
- What evidence would cause us to change or withdraw our approval?
School leaders are in a position to make a valuable contribution to AI safety by making AI adoption deliberate, visible and collective.
Further Guidance
- Department for Education (2026), Generative AI: product safety standards. Last updated 19 January 2026.
- Department for Education (2025), Generative artificial intelligence (AI) in education. Last updated 12 August 2025.