The Artificial PostAccount
All papers
Efficient AI · ORIGINAL RESEARCH

EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL

Selected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Jimmy Ba