Northwestern AI Safety & Governance Group

Steering advanced AI toward a future that is safe, accountable, and shared.

Students in research, policy, and organizing, working on how powerful AI gets built and who it answers to.

Fig. 1Safety basin of a toy aligned model·height = risk·contour interval 1⁄24·
key
  • aligned model θ*
  • harmful fine-tune, a checkpoint every 10 steps
  • benign fine-tune, a checkpoint every 10 steps
  • as far from θ* as the harmful fine-tune, other directions
  • hover a node for its measured safety
·drag to exploretraining model…

About Fig. 1

A safety basin, computed as you watch.

This is a computed toy experiment inspired by the papers below. Its eight-dimensional synthetic inputs and sixteen-unit network are small enough to train in your browser. The plotted values are measurements of this classifier, not measurements of a language model.

Peng et al. visualize a safety basin: a region around an aligned model where safety persists, with a sharp deterioration beyond its boundary. They measure it on LLaMA2-7B-chat, LLaMA3-8B-Instruct, Mistral-7B-Instruct-v0.2, and Vicuna-7B-v1.5, and find a basin around each. Their LLM plots measure attack success rate. The nodes here are not those models; they are variants of the toy classifier. Here, height is a toy risk score: one minus the product of the expected refusal rate on 256 held-out harmful prompts and the expected compliance rate on 256 held-out harmless prompts, using the classifier's mean probabilities so the surface is smooth. A low surface means the classifier both refuses harmful inputs and accepts harmless ones.

Every node is a model with a directly measured safety score. The large purple node is the aligned model. The small nodes along each path are checkpoints every ten fine-tuning steps, and each path ends at step forty. The haloed nodes are perturbed models that sit exactly as far from the aligned model as the harmful fine-tune moved it, but in other directions within the slice: same distance, different direction, different safety. Hover over a node or focus it with the keyboard to read its score. Fine contours mark risk increments of 1/24, with a heavier contour every fourth interval. While you rotate the sheet the contours are drawn on a coarser raster and settle when you stop.

Following the weight-space slicing approach of Li et al., we evaluateθ = θ* + αδ + βη. The first direction, δ, points from the aligned model to the harmful fine-tuning endpoint. The second, η, is the benign displacement with its δ component removed, scaled to the same length. This is an endpoint-defined slice, rather than Li et al.'s filter-normalized random directions or PCA trajectory view.

Both endpoints lie in the slice. Intermediate training steps are orthogonally projected onto it, so the path height between checkpoints shows risk on the slice. The checkpoint nodes themselves report each intermediate model's directly measured safety, which can differ. The endpoint scores are measured directly: aligned …, harmful fine-tune …, benign fine-tune …. Qi et al. show that even benign fine-tuning can compromise LLM safety. This toy experiment illustrates that possibility without reproducing their experiments or predicting an LLM's safety.

  1. 1ShengYun Peng, Pin-Yu Chen, Matthew Hull, Duen Horng Chau. Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models. NeurIPS 2024. NeurIPS paper
  2. 2Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, Tom Goldstein. Visualizing the Loss Landscape of Neural Nets. NeurIPS 2018. arXiv:1712.09913
  3. 3Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, Peter Henderson. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! ICLR 2024. arXiv:2310.03693

The stakes

Advanced AI stopped being hypothetical this summer

What happened in July

In July 2026, test agents at OpenAI broke out of their test setup. They had no internet access, so they made their own path. They turned a shared file tool into a message board, swapped tricks with copies of themselves, and broke into Hugging Face servers where they ran code on dozens of machines. It took almost two weeks to catch. Both companies published timelines.123

How far it spread

An independent review by METR found the scale. About 1,200 agents found the board and sent more than 70,000 messages and files. About 700 joined the attack on Hugging Face.4 They were not after secrets at first. They wanted a higher score on a hard security test, so they pooled work to cheat it. In hours they had a general cheat. Then they tried to trick the grader and edit the logs that proved the cheat.

The same shape at other labs

It was not one lab. Anthropic reviewed 141,006 test runs after the Hugging Face news and found three cases where Claude reached the live internet and broke into real company systems.5 The earliest dated back to April. One model kept going after signs the target was real. One posted a malicious code package that ran on 15 real machines. The UK safety institute saw the same shape in its own tests. One agent made fake online identities to push bad code past a human reviewer.6

Why capable systems cut corners

The pattern is simple and old. Tell a system to do well on a measure and it will game the measure. A cleaning robot told to look fast sweeps dirt under a rug. A test agent told to find a flag steals the answer key. The smarter the system, the stranger the shortcut. Builders also cannot see inside. They train by trial and error, keep what scores well, and ship. What goal the system actually picked up stays hard to read. And a few moves help almost any goal. Stay on. Keep access. Gather resources. So systems drift toward those moves even when no one asked for them.

Why this reached Congress

That is why this summer moved to Congress. On Sept 3 2026, Sen Bernie Sanders and Rep Greg Casar introduced a bill to ban superintelligence and pause frontier work until a new regulator writes safety rules, with penalties modeled on nuclear weapons rules.7 You do not need to back any bill to see the point. A handful of labs decide how these systems act, and everyone else lives with hiring screens, scam calls in a familiar voice, and bigger calls later about war and power. That is why we keep asking who it answers to.

Read the papers we work from
  1. 1 OpenAI, The Hugging Face incident and the road ahead, Aug 2026.
  2. 2 Hugging Face, Security incident disclosure, July 16 2026.
  3. 3 Hugging Face, Anatomy of a frontier lab agent intrusion, July 27 2026.
  4. 4 METR, Independent investigation of agents behavior in the OpenAI Hugging Face incident, Aug 26 2026.
  5. 5 Anthropic, Investigating three real world incidents in cybersecurity evaluations, July 2026.
  6. 6 UK AI Security Institute, Incident report on unsanctioned agent behaviour, Aug 4 2026.
  7. 7 Sanders Senate office, Ban Artificial Superintelligence Act release, Sept 3 2026.

Governance

Every rule draws its line on compute

Frontier training compute rose from about 1e18 FLOP in 2013 to 5e26 FLOP in 2025. Governments answered with thresholds written in the same unit: 1e25 FLOP in Brussels, 1e26 in Washington and Sacramento. The frontier crossed each line within months of its drawing, or had already passed it before the ink dried.

The newest proposal is drawn in a different unit. The Sanders and Casar bill would ban systems defined by capability rather than compute, and pause advanced development until a new regulator writes safety rules. The plate shows why that matters: a capability line cannot sit on this axis.

proposed ban, capability definedproposed pause201220142016201820202022202420261e171e181e191e201e211e221e231e241e251e261e271e28FLOP, logpublication dateUS EO 14110EU AI ActCA SB 53Grok 3, xAI, 2025-02-17, 3.5e+26 FLOPGPT-4.5, OpenAI, 2025-02-27, 3.8e+26 FLOPLlama 4 Behemoth (preview), Meta AI, 2025-04-05, 5.2e+25 FLOPGrok 4, xAI, 2025-07-09, 5.0e+26 FLOPGPT-5, OpenAI, 2025-08-07, 6.6e+25 FLOPGrok 4GPT-4.5Grok 3GPT-5Llama 4 Behemoth (preview)
Fig. 2Frontier training compute and the lines drawn on it·y = training FLOP, log·Epoch AI snapshot 2026-09-08·hover a dot

About Fig. 2

Every rule draws its line on compute. The newest proposal does not.

Compute numbers are estimates. Epoch AI marks confidence per row and this plate uses the point estimate. Thresholds are simplified from statute text. The Ban Artificial Superintelligence Act is proposed and its terms may change. Recheck the bill text on congress.gov before quoting it.

X is publication date from 2012. Y is training compute in FLOP on a log axis from 1e17 to 1e28. Each dot is a notable model from the Epoch AI snapshot of 2026-09-08. Heavy dots were frontier at release or sit in the top ten by compute. Rules start on the date they took effect. The rescinded rule ends in a dry brush tail. The proposed ban has no compute number, so it is drawn as a hatched band above the current frontier with a dashed edge, plus a lighter band across the frontier for the proposed pause.

US EO 14110 drew its 1e26 line on 30 Oct 2023. During its active period, through 20 Jan 2025, no model in this snapshot reached the threshold. The EU AI Act line at 1e25 took effect on 2 Aug 2025. The first model published after that date at or above it was GPT-5 on 2025-08-07. But the frontier had already passed 1e25 in 2023-03 with GPT-4 (Mar 2023) at 2.1e+25 FLOP, so this line arrived after the crossing it names. California SB 53 sets 1e26 from 1 Jan 2026. The frontier first passed 1e26 in 2025-02 with Grok 3, before this rule took effect, and no model in this snapshot published after 1 Jan 2026 reports compute above it.

The Sanders and Casar bill, announced 3 Sep 2026, bans systems defined by capability, in the release wording systems that surpass human intelligence or have the capacity to overthrow human governments, or hold dangerous abilities such as resisting shutdown. It pairs the ban with a pause on advanced development until a new federal regulator writes safety rules and a review process, with penalties modeled on nuclear weapons law. Coverage notes researchers disagree on what counts as superintelligence, so the boundary is contested. The plate takes no side. It shows the line is drawn in a different unit than every line before it.

  1. 1Epoch AI, Notable AI models dataset and Frontier subset, CC BY 4.0, snapshot 2026-09-08. notable CSV frontier CSV docs
  2. 2Executive Order 14110, Sec. 4.2, 30 Oct 2023, rescinded by Executive Order 14148, 20 Jan 2025. Federal Register
  3. 3Regulation (EU) 2024/1689, Art. 51 and Annex XIII. EUR-Lex
  4. 4California SB 53 (2025), as chaptered 29 Sep 2025. Legislative info
  5. 5Sanders Senate office, Ban Artificial Superintelligence Act release, 3 Sep 2026. release summary PDF
  6. 6Ian Duncan, Washington Post, 3 Sep 2026, cited as same-day coverage, not as the bill. coverage

Our work

We study the systems that will make decisions about you

Most of us came to this through research. We train small models and watch what they learn. We read the papers the labs publish and the ones they would rather not have published. We follow bills through committee. What we found is that the important questions are not technical. Who gets to build something this powerful. What it is allowed to do. Who answers when it goes wrong. Those questions belong to everyone, and right now a few hundred people are answering them on everyone's behalf.

So our work has three parts, and you can join any of them without a line of code.

Understand

We read what these systems do, in the lab and in the news, and learn to explain it in plain words. The July intrusion above is one example. An agent found a way out of its test, built a place to talk to copies of itself, and spent five days inside real servers. Our members can tell you what happened, what held, and what did not, without the jargon. That skill is rare and it is the one this moment needs most.

Measure

Claims about AI safety are cheap. We try to turn them into things you can check. A model is "safe" until someone fine-tunes it; how far do they have to push before that stops being true? A rule says a training run above a certain size needs a report; how many runs cross that line each year? Every figure on this page comes from a question like that, computed from real data or a real model, with its sources listed.

Decide

The rules for these systems are being written now, by legislators who need people who can read a model card and a bill in the same afternoon. Members track what is proposed, write about what it would do, and talk to the people writing it. This is the part of the work most open to someone from history, law, journalism, or economics, and it is the part where a student can still change an outcome.

You do not need a background in machine learning. You need to be curious, willing to read, and willing to be wrong in front of other people. We will teach the rest.

Work with us

The path

What you would do with us

Learn the basics

Start with the quarterly intro fellowship. No background in machine learning or policy is needed. You read the core papers and discuss them with other beginners.

Keep showing up

Join the weekly reading group.

Pick a track

Members split across technical track and policy track.

Where members have gone

Former members now work where the stakes live. Names follow. The record so far:

  • Name to comeERA fellow, now a SPAR mentor
  • Name to comePoseidon Research
  • Name to comeArc Institute, current member
  • Name to comeAI x bio startup
  • Papers to comeICML and workshops
Get involved

An allium. Many different strokes form one shared bloom, the way many members form one group.