The Artificial PostAccount
All papers
Language models · ORIGINAL RESEARCH

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Selected research in language models. Read the full paper, including the methods, experiments, and reported results.

Opening the paper…

KEEP FOLLOWING THE IDEA

Meet the researchers.

Danqi Chen