Geoffrey Hinton
University of Toronto
Learning useful internal representations from data.
Selected work Distilling the Knowledge in a Neural Network50 researchers
University of Toronto
Learning useful internal representations from data.
Selected work Distilling the Knowledge in a Neural NetworkUniversité de Montréal
Learning distributed representations that can support generalization.
Selected work Neural Machine Translation by Jointly Learning to Align and TranslateNew York University
Learning useful visual features and predictive representations.
Selected work Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureLanguage models
Using large neural networks to learn sequence transformations.
Selected work Sequence to Sequence Learning with Neural NetworksGenerative models
Training a generator through competition with a discriminator.
Selected work Generative Adversarial NetworksStanford University
Connecting visual recognition, language, and embodied intelligence.
Selected work ImageNet Large Scale Visual Recognition ChallengeStanford University
Learning features from unlabeled data at scale.
Selected work Building high-level features using large scale unsupervised learningAI for science
Combining learning and search to solve structured problems.
Selected work Gemma 4 Technical ReportReinforcement learning
Learning decisions through interaction and search.
Selected work Playing Atari with Deep Reinforcement LearningUniversity of Alberta
Learning predictions and actions from reward and experience.
Selected work True Online Temporal-Difference LearningProbabilistic learning
Using probabilistic structure to learn from complex data.
Selected work Peptide-Spectra Matching from Weak SupervisionGenerative models
Learning latent-variable models with efficient gradient estimators.
Selected work Auto-Encoding Variational BayesUniversity of Amsterdam
Using probabilistic latent spaces to model observations.
Selected work Auto-Encoding Variational BayesUniversity of Toronto
Making optimization and attention more efficient.
Selected work Adam: A Method for Stochastic OptimizationComputer vision
Connecting neural language generation with visual understanding.
Selected work Deep Visual-Semantic Alignments for Generating Image DescriptionsEfficient AI
Building the systems that make large-scale learning practical.
Selected work TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed SystemsLanguage models
Learning sequence representations and scaling neural networks.
Selected work Sequence to Sequence Learning with Neural NetworksLanguage models
Treating structured predictions as sequences.
Selected work Sequence to Sequence Learning with Neural NetworksAI safety
Studying scaling and the reliability of learned systems.
Selected work Scaling Laws for Neural Language ModelsScaling laws
Measuring how loss changes with model size, data, and compute.
Selected work Scaling Laws for Neural Language ModelsScaling laws
Studying predictable relationships between resources and learning.
Selected work Scaling Laws for Neural Language ModelsMultimodal AI
Learning transferable representations through large-scale pretraining.
Selected work Learning Transferable Visual Models From Natural Language SupervisionReinforcement learning
Improving policies while controlling the size of each update.
Selected work Proximal Policy Optimization AlgorithmsAI safety
Studying helpfulness, honesty, and harmlessness in AI assistants.
Selected work A General Language Assistant as a Laboratory for AlignmentInterpretability
Investigating the internal features and computations of networks.
Selected work Toy Models of SuperpositionLanguage models
Combining attention and sparse conditional computation.
Selected work Attention Is All You NeedLanguage models
Using attention to process sequences in parallel.
Selected work Attention Is All You NeedStanford University
Connecting linguistic structure with learned language representations.
Selected work Direct Preference Optimization: Your Language Model is Secretly a Reward ModelStanford University
Making language-model evaluation broader and more transparent.
Selected work Holistic Evaluation of Language ModelsStanford University
Learning adaptable skills from limited experience.
Selected work Model-Agnostic Meta-Learning for Fast Adaptation of Deep NetworksUC Berkeley
Learning policies for difficult perception and control problems.
Selected work Denoising Diffusion Probabilistic ModelsUC Berkeley
Learning robot behavior from data with deep reinforcement learning.
Selected work Model-Agnostic Meta-Learning for Fast Adaptation of Deep NetworksMIT
Making deep visual models easier to optimize and transfer.
Selected work Deep Residual Learning for Image RecognitionComputer vision
Studying deep architectures for visual and language modeling.
Selected work Very Deep Convolutional Networks for Large-Scale Image RecognitionGenerative models
Generating data by learning to reverse a noising process.
Selected work Denoising Diffusion Probabilistic ModelsGenerative models
Learning score functions to guide generative sampling.
Selected work Generative Modeling by Estimating Gradients of the Data DistributionStanford University
Connecting probabilistic modeling, optimization, and generation.
Selected work Generative Modeling by Estimating Gradients of the Data DistributionPrinceton University
Reducing memory traffic and computation in sequence models.
Selected work FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessStanford University
Making learning systems more efficient and usable.
Selected work FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessCarnegie Mellon University
Modeling long sequences with structured state spaces.
Selected work Mamba: Linear-Time Sequence Modeling with Selective State SpacesMIT
Compressing and accelerating neural networks for deployment.
Selected work Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman CodingEfficient AI
Reducing the memory cost of large-model training and inference.
Selected work QLoRA: Efficient Finetuning of Quantized LLMsUniversity of Washington
Learning language systems that can be adapted efficiently.
Selected work QLoRA: Efficient Finetuning of Quantized LLMsPrinceton University
Combining retrieval with language understanding.
Selected work Dense Passage Retrieval for Open-Domain Question AnsweringReasoning
Eliciting intermediate reasoning from language models.
Selected work Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsLanguage models
Studying compute-efficient training and sparse language models.
Selected work Training Compute-Optimal Large Language ModelsMcGill University
Studying reproducibility and evaluation in machine learning.
Selected work Deep Reinforcement Learning that MattersStanford University
Studying commonsense knowledge and model behavior.
Selected work ATOMIC: An Atlas of Machine Commonsense for If-Then ReasoningCaltech
Learning operators and representations for scientific problems.
Selected work Fourier Neural Operator for Parametric Partial Differential EquationsGenerative models
Controlling image synthesis through style-based generator architectures.
Selected work A Style-Based Generator Architecture for Generative Adversarial Networks