
Self-supervised representation steering and weak-to-strong oversight: alignment tools that frontier labs will actually use.
AISANZ Fellows are people in Australia and New Zealand doing AI safety work we're excited about: technical, policy, and research work we'd happily introduce to anyone in the field. Fellows are appointed by invitation.

Self-supervised representation steering and weak-to-strong oversight: alignment tools that frontier labs will actually use.

Localising safety-relevant concepts in models during training to make them inherently interpretable.
Look out for AI Safety Australia & New Zealand at conferences and workshops. Our Fellows list us as an affiliation on their papers and talks.
If you're doing AI safety work in Australia or New Zealand, we'd like to hear about it. Tell us what you're working on.
Get in touch→