Proximal Policy Optimization Algorithms
Proximal Policy Optimization proposes a policy-gradient objective that limits overly large updates while simplifying implementation.
Paper & contextImproving policies while controlling the size of each update.
Schulman studies reinforcement learning and how to train useful language-model behavior. This selection links trust-region and proximal policy optimization with human-feedback training, reward-model limitations, and evaluation of model responses.
10 papers
Proximal Policy Optimization proposes a policy-gradient objective that limits overly large updates while simplifying implementation.
Paper & contextSelected research in reinforcement learning, robotics. Read the full paper, including the methods, experiments, and reported results.
Paper & contextInstructGPT studies fine-tuning language models using demonstrations and human preference judgments to better follow instructions.
Paper & contextSelected research in reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextThe GPT-4 technical report documents capabilities, evaluations and limitations of a multimodal model, while withholding many architecture and training details.
Paper & contextSelected research in reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in scaling laws, reinforcement learning. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources