Honours thesis · in progress · submitting November 2026
What I’m working on
Can a sparse autoencoder find the exact part of a language model that holds a stereotype — precisely enough to cut just that part out?
The usual fix for bias in a language model is a broad edit, and a broad edit costs the model's general ability along with the bias. This thesis tests a narrower one: whether sparse autoencoders can find the features behind a single bias category — like disability or socioeconomic status — cleanly enough that cutting only those costs less. The method is built so a no counts as an answer.
Three stages. First, isolate candidate features by contrasting matched prompt pairs that differ only in the demographic term. Second, check what those features actually respond to, against a permutation-tested null. Third, cut them and weigh the drop in bias against the cost to general language ability — with a reconstruction-only control, so that cost isn't confused with the SAE's own reconstruction error.
What success looks like
A model that can have its disability-stereotype features switched off without also losing its grip on grammar or arithmetic or anything else — measured against a reconstruction-only control, not assumed. If cutting the bias costs about as much as a blunt, model-wide edit would, the honest answer is that this approach doesn't win, and the thesis says so.
University of Sydney · supervised by Dr. Huaming Chen
Mechanistic interpretabilitySparse autoencodersAI fairnessGPT-2TransformerLensSAELens