A General Language Assistant as a Laboratory for Alignment
Selected research in ai safety, scaling laws. Read the full paper, including the methods, experiments, and reported results.
Paper & contextStudying helpfulness, honesty, and harmlessness in AI assistants.
Askell’s coauthored research examines how AI assistants should behave and how that behavior can be evaluated. These papers investigate helpfulness, honesty, harmlessness, principles for alignment, and failures that can survive conventional safety training.
10 papers
Selected research in ai safety, scaling laws. Read the full paper, including the methods, experiments, and reported results.
Paper & contextConstitutional AI explores training a more helpful and harmless assistant with written principles and AI-generated feedback.
Paper & contextInstructGPT studies fine-tuning language models using demonstrations and human preference judgments to better follow instructions.
Paper & contextGPT-3 investigates how a large language model can perform tasks from instructions and examples in its prompt, without task-specific weight updates.
Paper & contextCLIP learns visual concepts from paired images and text, connecting language descriptions to image representations.
Paper & contextBIG-bench collects a broad set of language-model evaluation tasks to study capabilities and limitations beyond a single benchmark.
Paper & contextSelected research in scaling laws, ai safety. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in scaling laws, ai safety. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in scaling laws, ai safety. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in scaling laws, ai safety. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources