A starting point.
Adam combines adaptive learning rates with moving averages of gradients, becoming a widely used optimizer for neural networks.
What to keep in mind
Read the original paper for the experimental setting, baselines, and limitations; a result in one setting is not a guarantee of performance elsewhere.
This work is included in a researcher’s reading path. A detailed editorial explanation is still being prepared. The complete manuscript is available in Full paper.
Source: Adam: A Method for Stochastic Optimization. The original manuscript contains the methods, experiments, figures, and references. An arXiv posting date may follow an earlier conference publication. Read the linked record for version history.