School Connection / Flagship analysis
AI in schools: the evidence test for buy, trial or stop
The useful question is not whether a product contains AI, but whether a defined use can pass five gates: purpose, legal and safeguarding control, product assurance, a trial allowed to fail, and a credible route to scale or exit.
The answer in 60 seconds
Deploy a narrow use under controls. Trial uncertain benefit. Stop at every unresolved red line.
The strongest near-term case is low-risk, reversible, teacher-facing assistance in which no personal or confidential data is entered and a competent professional checks every output. Pupil-facing, data-bearing, analytical, safeguarding or agentic uses need stronger evidence because mistakes can affect children, rights, assessment or critical systems.
- Approve a use, not “AI”. Record the exact workflow, service plan, model/version, data, owner and fallback.
- A trial must be allowed to fail. Decide the comparator, non-inferiority thresholds and stop rules before launch.
- No source cited here certifies a product. Safety standards, framework routes, adoption and testimonials are not impact evidence.
01 / Controlled deploy, trial or stop
“Buy” is the end of a decision process, not a category of technology.
DfE says teacher-facing generative AI offers more immediate benefits and fewer risks than pupil-facing use, while the professional and organisation remain responsible. It also says evidence remains limited on learner development, educational outcomes and child safety.1 DfE’s 13-area safety standard widens the test beyond privacy and filtering to cognition, relationships, mental health and manipulation.2
Bounded staff drafting, formatting, translation or first-pass resources without personal or confidential data; low-stakes reversible administration.
Approved institutional account, acceptable data and IP terms, no autonomous consequential action, professional review, named owner, version record and non-AI fallback.
Pupil tutors and chatbots, data-bearing staff tools, feedback or marking support, attendance or behaviour analytics, safeguarding monitoring and agents that can act.
Time-limited and supervised, with baseline, comparator, DPIA screening, subgroup and accessibility checks, incident log, full cost and pre-agreed stop rules.
Opaque data chains, forced secondary use, absent child-safety controls, token oversight, AI as the sole marker in an assessment that forms part of an Ofqual-regulated qualification, unrestricted agents, inaccessible systems or no exit.
Do not connect live data, expose pupils, purchase or scale while any legal, safeguarding, equality, assessment, cyber, evidence or exit red line remains.
The standards are a non-statutory benchmark, not certification, an approved-product register or proof of learning impact.2 Ofsted neither requires AI nor evaluates products as a stand-alone inspection item; it says the evidence is still insufficient to define “good” use.10
02 / Why 2026 is different
An old AI checklist no longer covers the current decision.
- Product safety widened: DfE’s January 2026 update expanded its mainly supplier-facing product safety guidance to 13 areas, including cognitive development, emotional and social development, mental health and manipulation. DfE says schools and colleges may use the standards as a product-safety benchmark; they are not certification or proof of impact.2
- Data law changed: all data-protection provisions of the Data (Use and Access) Act 2025 were in force by 19 June 2026. It amends rather than replaces UK GDPR and the Data Protection Act. DUAA opened the full range of lawful bases for significant automated decisions where appropriate safeguards continue to apply; special-category data remains subject to stricter conditions.7
- Provider evidence matters: ICO reported 28 consensual edtech audits and 596 recommendations, with recurring gaps in roles, contracts, data maps, minimisation, retention, privacy and DPIAs. This is not a prevalence estimate.8
- Assessment boundaries sharpened: Ofqual expects awarding organisations to demonstrate valid, reliable and fair outcomes, with appropriate expert human involvement, quality assurance and accountability; it says AI may not be the sole marker in any assessment that forms part of a regulated qualification.11 Its January 2026 working paper adds that validity evidence must fit the assessment context and that agreement with human marks alone is insufficient.22
- Agentic systems reached the risk register: NCSC calls for bounded pilots, least privilege, temporary credentials, logging, sandboxing, active oversight and a stop mechanism. Its August 2026 advice is interim.199
- School standards became more AI-specific: DfE updated filtering, cyber, governance and accessibility standards on 25 August 2026.5161718
- Leadership support was refreshed: DfE updated its school and college leader materials for 2026/27 on 19 May. They support audit, strategy and implementation; they are not effectiveness evidence or product approval.24
At this article’s 29 August 2026 evidence cut-off, KCSIE 2025 remained in force through 31 August and KCSIE 2026 took effect on 1 September. Decisions and later updates must cite the version then in force.15
03 / The five-gate decision framework
Any red gate means stop. Amber means remediation or a bounded trial, not silent acceptance.
The five gates are School Connection’s synthesis of the cited law, statutory guidance, regulator positions, non-statutory standards and official evidence. They are not a DfE or regulator-issued approval scheme.
Problem, people and proportionality
Test: What exact problem, baseline and affected group justify AI over a simpler, lower-risk option?
Decision record: Outcome, stakes, users, measurable benefit, exact service plan and model/version, owners and fallback.
Legal, safeguarding and assessment hard stops
Test: Can the DPO, DSL and relevant curriculum or exams lead evidence every applicable duty before live use?
Decision record: Lawful basis, minimisation, DPIA screening, a completed DPIA where required, profiling/ADM, KCSIE, filtering, age, IP, equality, assessment and procurement route.
Product and supplier assurance
Test: Do current documents, tests and demonstrations support this use, rather than merely the product category?
Decision record: Safety case, realistic testing, security, data flow, child controls, accessibility, change control, evidence, cost and exit.
A bounded pilot that can fail
Test: Will the pilot produce a decision, including a credible decision to stop?
Decision record: Users, duration, baseline, comparator, primary benefit, non-inferiority thresholds, subgroups, incidents, full cost and stop rules.
Scale, monitor, renew or exit
Test: Has the exact use shown net benefit without safety, quality or equality deterioration?
Decision record: Approved use/version register, restrictions, KPIs, change notice, audit, export, deletion, suspension, fallback and review date.
DfE’s current school data and EdTech guidance supports this whole-chain view: lawful basis, minimisation, roles, sub-processors, contracts, security, retention, deletion and exit all sit inside the decision.34 For academy trusts, a framework can streamline procurement, but it does not remove the trust’s responsibilities as contracting authority. Neither a framework nor an approved buying option establishes AI safety, data-law compliance, accessibility, educational effectiveness or value for the intended use.25
Equality is not a supplier checkbox. Relevant public authorities in Great Britain must exercise due regard under the Public Sector Equality Duty before and during the decision and continue to monitor effects; the duty cannot be delegated to a vendor or an impact-assessment document.21 DfE’s accessibility standard also places accessibility evidence, assistive-technology compatibility and accessible procurement inside the decision.18
04 / The pupil-facing test
Child safety is more than filtering harmful words.
Under this article’s decision framework, pupil-facing tutors, chatbots, adaptive practice and writing aids should default to supervised, time-limited trials in 2026. DfE requires close supervision and appropriate age, safety, filtering and monitoring controls; the “trial” classification is School Connection’s evidence-led recommendation, not a statutory rule. The governance conclusion is strong; general product-effectiveness evidence is not.
- Age-appropriate filtering, usable logs and a named local safeguarding alert route
- Close supervision, age compliance and controls that work for SEND and EAL contexts
- Progressive disclosure and prompts for the pupil to attempt work before receiving an answer
- Measures of cognitive offloading rather than completion alone
- No implied personhood, emotional dependence, sycophancy or manipulation
- Human routes for distress, session limits and no commercial steering
These expectations come directly from DfE’s education policy, product standards and filtering-and-monitoring standard.125 They are reasons to test carefully, not evidence that a product meeting them improves learning.
Do not proceed where a pupil bot implies personhood, seeks emotional disclosure, cultivates dependence, extends engagement, behaves sycophantically or steers the child commercially. Do not proceed where complete answers are the default and progressive disclosure cannot be enforced.
05 / The data map
If the organisation cannot draw the chain, it cannot approve the use.
What content and metadata enter the service?
Who is controller, joint controller or processor for each purpose?
Which model, API, plug-in, sub-processor and data location sits in the chain?
Are inputs or outputs used for training, product development, advertising or another secondary purpose?
What lawful basis, minimisation, retention, deletion and rights process applies?
Can the school export what it needs, verify deletion, including backups, and continue without the service?
A DPIA must be completed before processing that is likely to result in a high risk to individuals’ rights and freedoms; AI alone does not automatically trigger one. Screen every use and document the reasons where a DPIA is not required. Children, profiling, systematic monitoring, sensitive data and innovative technology can combine to trigger the requirement.6 For a decision not to count as solely automated, human involvement must be meaningful and active rather than token. Where a solely automated significant decision is permitted, safeguards include information, a way to express a view, the ability to obtain human intervention and a challenge route.20
Contract labels are not conclusive. ICO says the facts and purposes determine whether an edtech provider is a processor, controller or joint controller and whether the Children’s Code applies to its service.26
06 / Human oversight
“A human is in the loop” is not a control unless the human can change the result.
For a consequential output, the reviewer must be competent, see the relevant evidence, have enough time, understand the system’s limits, be able to reject or change the result and record a reason. Where UK GDPR articles 22A to 22D apply, safeguards include information, a way to express a view, human intervention and a challenge route.720
The same principle holds in assessment. Ofqual’s position requires expert human involvement and prevents AI acting as the sole marker in a regulated qualification assessment.1122 A low-stakes classroom feedback trial does not establish fitness for a high-stakes qualification.
The sole-marker boundary applies to assessments that form part of Ofqual-regulated qualifications; it is not a blanket prohibition on every low-stakes classroom feedback tool. AI-generated work submitted as a pupil’s own can be malpractice, so leaders should map coursework and non-exam-assessment exposure, set task-specific rules and follow current awarding-organisation and centre procedures when authenticity is in doubt.23
07 / A trial that can fail
Pre-register the decision before the excitement of the pilot changes the question.
Fix the users, site, subject, task, version and duration. Set a baseline and comparator. Choose one primary benefit and non-inferiority thresholds for quality, safety and equality. Measure net workload after checking and support; use blind quality review; test unaided learning; examine relevant subgroups; log incidents and near misses; count direct and hidden cost; and name the reviewer, decision date and stop rules before launch.
Net workload or learning
Measure minutes after checking, correction, training and support, or curriculum outcomes and unaided performance against a comparator.
Quality, safety and equality
Pre-agreed thresholds that prevent a time saving being bought with more errors, exclusion or safeguarding burden.
Incidents and stop rules
Harmful outputs, false and missed alerts, jailbreaks, material subgroup disadvantage, data incidents and contract or model changes.
Questions for leaders and trustees
What precise problem are we solving, and what is the current baseline?
Why is AI necessary rather than a simpler workflow, automation or policy change?
Who will use the tool, who will be affected and what is the highest-consequence plausible error?
What exact product, service plan, model/version, integration and data region are we approving?
Who can pause or stop use?
Does the use require a DPIA, and where is that screening decision recorded?
How do filtering, disclosures and local DSL escalation work in practice?
Can the human reviewer understand, reject and change a consequential output?
How has the use been tested with relevant ages, SEND, EAL and assistive technologies?
What checking, training, support, integration, moderation and exit cost sits outside the licence?
What supplier changes force the organisation to reopen approval?
What pilot result would make us stop, and on what date will the evidence be reviewed?
08 / Evidence boundary
Official evidence supports disciplined experimentation, not a general claim that AI improves learning.
82% of primary and 78% of secondary teacher respondents reported using generative AI in their role.14
Self-reported practice does not measure attainment, net workload, safety or value.
Ofsted’s early-adopter research reported leader perceptions of workload benefit.12
It interviewed 21 purposively selected leaders, no teachers or pupils, and was not an impact evaluation.
A DfE prototype detected 43 of 47 inserted errors in a constrained individual-sentence Year 4 literacy dataset.13
It did not test attainment, net teacher workload, long-term safety or fairness and cannot be generalised to products.
- A product name can conceal changing models, service plans, dependencies and data terms.
- Long-term evidence on children’s cognitive, relational and wellbeing outcomes remains sparse.
- Average results can conceal differential error, access and benefit for SEND, EAL and protected groups.
- Recorded malpractice sanctions do not estimate undetected AI misuse.23
- The ICO DPIA page cited here was under DUAA review at the evidence cut-off and must be rechecked before a later update.6
- NCSC’s August 2026 agentic advice is interim pending formal guidance.9
- Education-specific conclusions here apply to England; devolved regimes require separate review.
09 / What we are monitoring next
AI approval is not permanent.
Reopen the gate record after a material model, feature, service-plan, supplier-term, integration, data-category, user-group or autonomous-action change; after an incident, repeatable jailbreak, accessibility complaint or subgroup disadvantage; when law, KCSIE, regulator or awarding-organisation rules change; when NCSC replaces its interim agentic guidance; and at contract renewal or a claimed impact milestone.
The safest question is not “Can this tool produce a good answer?” It is “Can our organisation explain and control the whole decision, from purpose and data to pupil experience, human judgement, evidence, contract and exit?”
Editorial rebuild note
A ground-up evidence rebuild, not a cosmetic update.
This analysis replaces an article first published on 30 April 2026 as “AI in Schools: What Should Leaders Buy, Trial or Avoid in 2026?”. On 29 August 2026, School Connection removed the product suggestions, supplier framing and every legacy image; rechecked the material claims against primary DfE, ICO, NCSC, Ofsted and Ofqual sources; and separated binding duties from statutory guidance, regulator interpretation, non-statutory standards, research and policy. It does not endorse or certify any commercial product and is not legal advice.
Approved public intelligence
What the live evidence is showing now.
School Connection continues to monitor digital, data, ai & cyber resilience evidence. New machine-detected signals remain private editorial candidates until a human editor investigates and approves them for publication.
This panel reads only the editor-approved School Connection public feed. It never exposes raw Observatory records, private candidates, contacts, commercial opportunities or internal scores.
Sources and methodology
Evidence used in this analysis
School Connection links to the primary source behind each material claim. Source status, period and limitations are stated so readers can reproduce the evidence trail.
- 01
Department for Education · Updated 12 August 2025
Generative artificial intelligence in education
England policy and non-statutory guidance, including the limited evidence base and the lower-risk case for bounded teacher-facing use. - 02
Department for Education · Updated 19 January 2026
Generative AI: product safety standards
Mainly supplier-facing, non-statutory guidance that schools may use as a safety benchmark: 13 areas spanning technical, child, cognitive, relational, mental-health and manipulation controls. - 03
Department for Education · Updated 9 July 2026
Generative AI and data protection in schools
Official school guidance on approved tools, personal data, pupil intellectual property, age restrictions and privacy information. - 04
Department for Education · Updated 9 July 2026
Procuring educational technology
The data-flow, contract, DPIA, sub-processor, retention, deletion and exit checks behind the five-gate model. - 05
Department for Education · Updated 25 August 2026
Filtering and monitoring: core standard
Places generative AI within the filtering and monitoring picture, including dynamic content, local roles and SEND/EAL context. - 06
Information Commissioner’s Office · Retrieved 29 August 2026
When do we need to do a DPIA?
UK regulator guidance: a DPIA is required for likely high-risk processing, not simply because a product uses AI. The page was under DUAA review. - 07
Information Commissioner’s Office · Updated 19 June 2026
The Data (Use and Access) Act 2025: what it means for organisations
Explains that the Act amends rather than replaces UK GDPR, broadens lawful bases for significant automated decisions subject to safeguards, and keeps stricter conditions for special-category data. - 08
Information Commissioner’s Office · Published 24 June 2026
Statement on EdTech examined
Reports 28 consensual provider audits and 596 recommendations; it is not an AI-product prevalence study. - 09
National Cyber Security Centre · Published 20 August 2026
Managing the cyber risk of agentic AI
Interim cross-sector advice on sandboxing, active oversight, audit and stop mechanisms for systems that can act. - 10
Ofsted · Updated 21 October 2025
How Ofsted looks at AI during inspection and regulation
Ofsted does not require AI use or provide a stand-alone product evaluation; leaders remain accountable for outcomes, safeguarding and assurance. - 11
Ofqual · Updated 16 July 2026
Approach to regulating AI in the qualifications sector
The evidence-led position for awarding organisations, including expert human oversight and the boundary against AI as sole marker in any assessment forming part of a regulated qualification. - 12
Ofsted · Published 27 June 2025
Insights from early adopters of AI in schools and FE
Qualitative implementation evidence from 21 purposively selected leaders, not representative or causal impact evidence. - 13
Department for Education · Published 28 August 2024; updated 17 October 2024
Generative AI in education: user research and technical report
A constrained proof of concept that illustrates why technical accuracy cannot be generalised to products, subjects or learning impact. - 14
Department for Education · Updated 25 June 2026
School and college voice: December 2025
Official self-reported adoption evidence. Use rates do not establish net workload saving, learning impact or safety. - 15
Department for Education · KCSIE 2025 in force to 31 August 2026; 2026 edition from 1 September 2026
Keeping children safe in education
The statutory safeguarding transition relevant to the 2026/27 school-year opening. - 16
Department for Education · Updated 25 August 2026
Digital leadership and governance: core standard
Named leadership, outcome-linked strategy, approved-app and information registers, contract management and continuity. - 17
Department for Education · Updated 25 August 2026
Cyber security: core standard
Risk, least privilege, logging, patching, approved apps, incident response and supplier assurance. - 18
Department for Education · Updated 25 August 2026
Digital accessibility
Accessibility evidence, assistive-technology compatibility and accessible procurement expectations. - 19
National Cyber Security Centre · Published 15 May 2026
Thinking carefully before adopting agentic AI
Official advice to begin with bounded, low-risk pilots and prevent unrestricted sensitive or critical access. - 20
Information Commissioner’s Office · Retrieved 29 August 2026
Profiling children and automated decisions
Current UK regulator guidance on active human involvement and information, intervention and challenge safeguards. - 21
Government Equalities Office / Equality Hub · Published 18 December 2023
Public Sector Equality Duty: guidance for public authorities
Great Britain general-duty guidance: consider effects before and during the decision and continue monitoring. - 22
Ofqual · Published 14 January 2026
Principles of AI use in marking
A working paper on validity, variability, bias, transparency and why human-mark agreement alone is insufficient. - 23
Ofqual · Published 9 March 2026
Senior leadership team briefing pack: AI and coursework integrity
Official leadership guidance on coursework exposure, authenticity conversations and existing malpractice procedures. - 24
Department for Education · Updated 19 May 2026
Using AI in education: support for school and college leaders
Implementation support for 2026/27; not effectiveness evidence or product approval. - 25
Department for Education · Published 15 July 2026
How to reduce procurement risk: good practice for academy trusts
Governance, value, conflicts, competition, records, due diligence and contract-management expectations for England academy trusts. - 26
Information Commissioner’s Office · Updated 30 May 2023
The Children’s Code and education technologies
Explains that factual control and purpose, rather than a contract label, determine provider role and Code scope.