Max McWhae

This site is live

2026-07-28

This site is where I work in public.

The short version of what I’m doing: I work on technical AI safety from Perth. My main project is SANDGLASS, a study of whether open-weight language models internally represent “I am being evaluated,” and whether deliberate underperformance triggered by that recognition — sandbagging — can be caught behaviourally or with simple activation probes. It runs on a single consumer GPU, and everything it produces will be open: the dataset, the harness, the trained models, and the write-up, including any null results.

I also founded the Perth AI Safety Meetup, which is relaunching this semester as a monthly series at UWA. If you’re in Perth and you care about any of this — or you’re just curious — come along.

What to expect here:

Everything is subscribable by RSS. More soon.