The Artificial PostAccount
All papers
Language models · ORIGINAL RESEARCH

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Selected research in language models. Read the full paper, including the methods, experiments, and reported results.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Yejin Choi