Ariana Azarbal
Hello! These days, I'm a researcher through the Anthropic Fellows Program. I have broad interests in scalable oversight, AI psychology, and AI welfare. I recently worked on mitigations for reward hacking and misgeneralization, as a MATS fellow.
I also study Math-CS at Brown, where I help lead BAIST. Lately, I've been enjoying yoga and creative writing.
Recent Research
- Recontextualization
- Inoculation Prompting
- Training a Reward Hacker Despite Perfect Labels
- Selective Generalization: Improving Capabilities while Maintaining Alignment
Writings
- Confusion around the term reward hacking March 2026