This site is live
This site is where I work in public.
The short version of what I’m doing: I work on technical AI safety from Perth. My main project is SANDGLASS, a study of whether open-weight language models internally represent “I am being evaluated,” and whether deliberate underperformance triggered by that recognition — sandbagging — can be caught behaviourally or with simple activation probes. It runs on a single consumer GPU, and everything it produces will be open: the dataset, the harness, the trained models, and the write-up, including any null results.
I also founded the Perth AI Safety Meetup, which is relaunching this semester as a monthly series at UWA. If you’re in Perth and you care about any of this — or you’re just curious — come along.
What to expect here:
- Explainers. Starting with sandbagging itself: what it is and why the entire evaluation enterprise depends on catching it.
- Devlogs. SANDGLASS progress notes, including a public go/no-go decision three weeks in. If something fails, I’ll write that up too.
- Practical notes. What AI safety research is actually feasible on one RTX 3090 in 2026, and what I learn running a meetup from zero.
Everything is subscribable by RSS. More soon.