Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Direct Preference Optimization derives a preference-learning objective that avoids training a separate reward model in the studied setup.
Paper & contextConnecting linguistic structure with learned language representations.
Manning studies language understanding and learned representations of linguistic structure. The selected work connects language-model training and evaluation with questions about syntax, generalization, and how models use language in broader systems.
10 papers
Direct Preference Optimization derives a preference-learning objective that avoids training a separate reward model in the studied setup.
Paper & contextSelected research in language models, ai evaluation. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in computer vision, language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextBIG-bench collects a broad set of language-model evaluation tasks to study capabilities and limitations beyond a single benchmark.
Paper & contextSelected research in language models, robotics. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources