Adam: A Method for Stochastic Optimization
Adam combines adaptive learning rates with moving averages of gradients, becoming a widely used optimizer for neural networks.
Paper & contextMaking optimization and attention more efficient.
Ba studies optimization, representation learning, and efficient neural networks. Adam and layer normalization are central starting points in this selection, followed by work on training, evaluating, and adapting language models.
10 papers
Adam combines adaptive learning rates with moving averages of gradients, becoming a widely used optimizer for neural networks.
Paper & contextLayer normalization normalizes a network's activations within an individual training example rather than across a batch.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources