The Artificial PostAccount
Researchers

Chris Olah

Interpretability

Investigating the internal features and computations of networks.

Olah studies the features and computations inside neural networks. The selected work connects interpretability with model behavior and alignment, including toy models of superposition and experiments aimed at understanding what language models represent.

Selected work

10 papers

An independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.

Selection & sources