The Artificial PostAccount
All papers
Interpretability · ORIGINAL RESEARCH

Auditing language models for hidden objectives

Selected research in interpretability. Read the full paper, including the methods, experiments, and reported results.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Chris Olah