The Artificial PostAccount
All papers
Generative models · ORIGINAL RESEARCH

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking

Selected research in generative models. Read the full paper, including the methods, experiments, and reported results.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Ian Goodfellow