The Artificial PostAccount
Researchers

John Schulman

Reinforcement learning

Improving policies while controlling the size of each update.

Schulman studies reinforcement learning and how to train useful language-model behavior. This selection links trust-region and proximal policy optimization with human-feedback training, reward-model limitations, and evaluation of model responses.

Selected work

10 papers

2016 · Reinforcement learning

OpenAI Gym

Selected research in reinforcement learning. Read the full paper, including the methods, experiments, and reported results.

Paper & context
2023 · LLMs

GPT-4 Technical Report

The GPT-4 technical report documents capabilities, evaluations and limitations of a multimodal model, while withholding many architecture and training details.

Paper & context

An independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.

Selection & sources