
Sovereign AI for Schools: Keeping Student Data Where It Belongs
- Schools became the leading edge of the sovereignty conversation by accident. They hold dense, permanent records on minors, and in 2024 and 2025 that data started leaving the building in bulk.
- The PowerSchool breach disclosed in December 2024 exposed roughly 62 million students and 9.5 million teachers through a single stolen credential. The company paid a $2.85 million ransom, and attackers extorted individual districts anyway (TechCrunch).
- The risk has moved into the vendor cloud. Verizon found third-party-involved breaches roughly doubled to about 30% of all breaches in 2025 (Verizon DBIR), while 82% of K-12 schools reported a cyber incident in an 18-month span (CIS).
- FERPA already implies the answer. Its school-official exception requires direct control, single-purpose use, and no redisclosure, which a multi-tenant model that trains on student data quietly violates.
- Sovereign AI for schools is not a luxury tier. It is data residency and a hard no-training rule written into procurement, so useful AI can run without shipping student records somewhere they should not go.

Why Education Is the Leading Edge
Start with the adoption curve, because it is steeper than the policy curve and the gap between them is where the trouble lives. Pew found that about 26% of U.S. teens had used ChatGPT for schoolwork by late 2024, double the share a year earlier (Pew Research). Tyton Partners put regular student use of generative AI at 44% in 2025, up from 27% in 2023. Meanwhile, a Common Sense Media analysis found that only about 5% of districts had a specific generative-AI policy. Read those three numbers together and you get the actual state of affairs: nearly half the students are using these tools, and roughly one district in twenty has written down what that is allowed to mean.
Now layer on what schools are holding while this happens. Not preferences and purchase histories, but Social Security numbers, health records, special-education designations, free-lunch status, custody arrangements. This is identity-theft-grade material on people too young to have a credit file, which makes it more valuable, not less, because the fraud can ripen quietly for years. The combination of dense permanent data and thin governance is exactly the condition under which a sovereignty problem becomes a sovereignty crisis. The schools did not volunteer for the leading edge. They were drafted by the contents of their own filing cabinets.

The Breach That Made the Argument
If you want the single event that turned "we should think about where student data lives" into "we need to think about this now," it is PowerSchool. In December 2024, an attacker used one stolen credential to reach a student-information system that serves around 60 million students across North America, and made off with the personal data of roughly 62 million students and 9.5 million teachers, Social Security numbers and health information included. PowerSchool paid a ransom of about $2.85 million in the belief that doing so would make the data disappear.
It did not. By May 2025, criminals were emailing individual districts in Canada and North Carolina with samples of the stolen records, extorting the schools directly even though the vendor had already paid (EdWeek). The perpetrator turned out to be a 19-year-old college student, later sentenced to four years in federal prison and about $14.1 million in restitution, which tells you something uncomfortable about the ratio between the effort required and the damage done. The lesson districts took away was not "pick a better vendor." It was structural: when your students' data sits inside someone else's multi-tenant platform, one stolen password is a district-wide event, and a paid ransom buys you nothing you can rely on. That is the same vendor-control problem I keep circling back to in three questions to ask before you feed data to an AI vendor, just with the volume turned all the way up.

What FERPA Actually Demands
People treat FERPA like a vague cloud of caution hanging over edtech. It is more specific than that, and the specifics happen to describe a sovereign architecture whether or not anyone using them has heard the word. Under the school-official exception, a district can share student records with a vendor without parental consent only if four things are true: the vendor performs a function the school would otherwise do itself, it has a legitimate educational interest in the data, it uses that data only for the assigned purpose, and it remains under the school's direct control regarding how the data is used and kept.
Hold a typical consumer AI product up against that list. A model that trains on whatever you feed it is not using the data solely for your assigned purpose. A platform that pools tenants and reserves broad rights to improve its service is not under your direct control in any meaningful sense. The moment student data becomes training fuel or crosses into the vendor's own purposes, the school-official exception stops applying and you are back to needing consent you almost certainly do not have. FERPA, in other words, is not asking districts to be cautious in spirit. It is asking them to be able to prove control, which is a procurement requirement dressed as a privacy law. This is the practical heart of privacy by design: the compliance is in the architecture, not the disclaimer.

Keeping Data on Home Turf
Sovereign AI is a serious-sounding phrase for a plain idea: keep the data where you can see it. In practice that lands on a spectrum, and districts do not all need the same point on it. At one end is genuinely on-premise inference, your own hardware running an open-weight model so student records never leave the building. In the middle is a state-hosted or regional private cloud, which is how a lot of public-sector sovereignty actually gets done, since few districts want to run a GPU cluster next to the boiler room. At the other end is a contractual arrangement with a commercial provider that pins data residency to a jurisdiction, forbids training on your data, and gives you deletion and audit rights you can actually exercise.
The good news for anyone dreading a capital expense is that the small, open models got good enough to make the local end of that spectrum realistic, which I wrote about in the context of why small models quietly won the enterprise. You do not need a frontier model to draft a parent newsletter, summarize an IEP meeting, or answer a routine policy question. You need a competent model that runs somewhere you control. The sovereignty conversation has spent years sounding like an argument for spending more, when the more honest version is that it is an argument for knowing where your data is, and the cost of that knowledge has been falling. The infrastructure question, as I argued when digital sovereignty became an infrastructure bill, is finally answerable without a moonshot budget.

A Procurement Checklist for Districts
If sovereignty is mostly a matter of control, then control is mostly a matter of what you wrote into the contract before anyone got excited about the demo. A few clauses do most of the work. Require, in writing, that the vendor will not train any model on your student data, full stop, no "anonymized" loopholes. Pin data residency to a named jurisdiction so the data cannot drift to wherever compute is cheapest this quarter. Demand data minimization, so the tool ingests the narrowest slice of records it needs rather than a convenient copy of everything. Set retention and deletion terms you can verify, and forbid cross-tenant use outright. Then make the vendor show you the access logs and audit trail, because direct control you cannot inspect is just trust with extra paperwork.
None of this requires a district to become a security shop, and it pairs naturally with the kind of plain-language internal rules districts are already adopting, like the red, yellow, green "stoplight" policies CoSN highlighted at its 2026 conference for telling teachers and students what AI use is permitted on which task. The point is to make the safe path the easy path, so a teacher reaches for an approved tool instead of quietly pasting a class roster into a consumer chatbot at 10pm. Govern the procurement and you have governed most of the risk, which is the whole argument behind an AI governance framework people actually follow: the framework only matters if it survives contact with a tired human being trying to get their grading done.

Frequently Asked Questions
What is sovereign AI for education?
Sovereign AI for education means deploying AI so that student data stays inside the institution's own control. The data remains within the district or state trust boundary, the model does not train on student records, and the school can demonstrate direct control over where the data lives and who can touch it. In practice that can mean on-premise hardware running an open-weight model, a state-hosted private cloud, or a commercial contract that forbids the vendor from moving or reusing the data.
How do schools protect student data when using AI?
By treating data residency and use restrictions as procurement requirements rather than afterthoughts. That means contracts that prohibit training on student records, keep data in a defined jurisdiction, minimize what is collected, set verifiable retention and deletion rules, and forbid cross-tenant use. It also means access controls, logging, and a short list of approved tools, so the safe option is the convenient one.
What does FERPA require for AI tools?
For a vendor's AI tool to handle student records without parental consent, FERPA's school-official exception requires that the tool performs a function the school would otherwise do, has a legitimate educational interest, uses the data only for that purpose, and stays under the district's direct control. An AI vendor that trains on or reuses student data for its own purposes generally falls outside that exception and would need consent.
Can schools use ChatGPT under FERPA?
A school can use a general AI assistant for tasks that involve no personally identifiable student records, or under an enterprise agreement that contractually bars the provider from training on or retaining the data and keeps the district in control. The trouble starts when someone pastes identifiable student information into a consumer chatbot with no such agreement behind it.
Why are schools a frequent target for data breaches?
Schools hold dense, long-lived records on minors, Social Security numbers, health and special-education data, family information, and they increasingly hold it inside shared vendor platforms. When one of those platforms is breached, the blast radius spans thousands of districts at once, which is precisely what made the PowerSchool incident so large.
References
- TechCrunch, PowerSchool paid a hacker's ransom, but now schools say they are being extorted (May 8, 2025)
- EdWeek, PowerSchool Paid a Hacker's Ransom. Now Cyber Criminals Are Threatening Schools (May 2025)
- Pew Research Center, About a quarter of U.S. teens have used ChatGPT for schoolwork, double the share in 2023 (January 15, 2025)
- Comparitech, Education Ransomware Roundup 2025
- Center for Internet Security, 2025 K-12 Cybersecurity Report (2025)
- Verizon, 2025 Data Breach Investigations Report (2025)
- U.S. Department of Education, Student Privacy Policy Office / FERPA guidance
- EdTech Magazine, CoSN 2026: How K-12 Districts Are Tackling Responsible AI Adoption (April 2026)
