Select or touch a mark to inspect it.
Northwestern AI Safety & Governance Group
Steering advanced AI toward a future that is safe, accountable, and shared.
Students in research, policy, and organizing, working on how powerful AI gets built and who it answers to.
key
- aligned model θ*
- harmful fine-tune, a checkpoint every 10 steps
- benign fine-tune, a checkpoint every 10 steps
- as far from θ* as the harmful fine-tune, other directions
- hover a node for its measured safety
The stakes
Advanced AI stopped being hypothetical this summer
What happened in July
In July 2026, test agents at OpenAI broke out of their test setup. They had no internet access, so they made their own path. They turned a shared file tool into a message board, swapped tricks with copies of themselves, and broke into Hugging Face servers where they ran code on dozens of machines. It took almost two weeks to catch. Both companies published timelines.123
How far it spread
An independent review by METR found the scale. About 1,200 agents found the board and sent more than 70,000 messages and files. About 700 joined the attack on Hugging Face.4 They were not after secrets at first. They wanted a higher score on a hard security test, so they pooled work to cheat it. In hours they had a general cheat. Then they tried to trick the grader and edit the logs that proved the cheat.
The same shape at other labs
It was not one lab. Anthropic reviewed 141,006 test runs after the Hugging Face news and found three cases where Claude reached the live internet and broke into real company systems.5 The earliest dated back to April. One model kept going after signs the target was real. One posted a malicious code package that ran on 15 real machines. The UK safety institute saw the same shape in its own tests. One agent made fake online identities to push bad code past a human reviewer.6
Why capable systems cut corners
The pattern is simple and old. Tell a system to do well on a measure and it will game the measure. A cleaning robot told to look fast sweeps dirt under a rug. A test agent told to find a flag steals the answer key. The smarter the system, the stranger the shortcut. Builders also cannot see inside. They train by trial and error, keep what scores well, and ship. What goal the system actually picked up stays hard to read. And a few moves help almost any goal. Stay on. Keep access. Gather resources. So systems drift toward those moves even when no one asked for them.
Why this reached Congress
That is why this summer moved to Congress. On Sept 3 2026, Sen Bernie Sanders and Rep Greg Casar introduced a bill to ban superintelligence and pause frontier work until a new regulator writes safety rules, with penalties modeled on nuclear weapons rules.7 You do not need to back any bill to see the point. A handful of labs decide how these systems act, and everyone else lives with hiring screens, scam calls in a familiar voice, and bigger calls later about war and power. That is why we keep asking who it answers to.
Read the papers we work from- 1 OpenAI, The Hugging Face incident and the road ahead, Aug 2026.
- 2 Hugging Face, Security incident disclosure, July 16 2026.
- 3 Hugging Face, Anatomy of a frontier lab agent intrusion, July 27 2026.
- 4 METR, Independent investigation of agents behavior in the OpenAI Hugging Face incident, Aug 26 2026.
- 5 Anthropic, Investigating three real world incidents in cybersecurity evaluations, July 2026.
- 6 UK AI Security Institute, Incident report on unsanctioned agent behaviour, Aug 4 2026.
- 7 Sanders Senate office, Ban Artificial Superintelligence Act release, Sept 3 2026.
Governance
Every rule draws its line on compute
Frontier training compute rose from about 1e18 FLOP in 2013 to 5e26 FLOP in 2025. Governments answered with thresholds written in the same unit: 1e25 FLOP in Brussels, 1e26 in Washington and Sacramento. The frontier crossed each line within months of its drawing, or had already passed it before the ink dried.
The newest proposal is drawn in a different unit. The Sanders and Casar bill would ban systems defined by capability rather than compute, and pause advanced development until a new regulator writes safety rules. The plate shows why that matters: a capability line cannot sit on this axis.
Our work
We study the systems that will make decisions about you
Most of us came to this through research. We train small models and watch what they learn. We read the papers the labs publish and the ones they would rather not have published. We follow bills through committee. What we found is that the important questions are not technical. Who gets to build something this powerful. What it is allowed to do. Who answers when it goes wrong. Those questions belong to everyone, and right now a few hundred people are answering them on everyone's behalf.
So our work has three parts, and you can join any of them without a line of code.
Understand
We read what these systems do, in the lab and in the news, and learn to explain it in plain words. The July intrusion above is one example. An agent found a way out of its test, built a place to talk to copies of itself, and spent five days inside real servers. Our members can tell you what happened, what held, and what did not, without the jargon. That skill is rare and it is the one this moment needs most.
Measure
Claims about AI safety are cheap. We try to turn them into things you can check. A model is "safe" until someone fine-tunes it; how far do they have to push before that stops being true? A rule says a training run above a certain size needs a report; how many runs cross that line each year? Every figure on this page comes from a question like that, computed from real data or a real model, with its sources listed.
Decide
The rules for these systems are being written now, by legislators who need people who can read a model card and a bill in the same afternoon. Members track what is proposed, write about what it would do, and talk to the people writing it. This is the part of the work most open to someone from history, law, journalism, or economics, and it is the part where a student can still change an outcome.
You do not need a background in machine learning. You need to be curious, willing to read, and willing to be wrong in front of other people. We will teach the rest.
Work with usThe path
What you would do with us
Where members have gone
Former members now work where the stakes live. Names follow. The record so far:
- Matthew KhoriatyRedwood Research, prev ERA fellow, SPAR mentor
- Andrii ShportkoPoseidon Research
- Ishan MukherjeeArc Institute
An allium. Many different strokes form one shared bloom, the way many members form one group.