Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Deep Compression combines pruning, weight quantization, and coding to reduce the storage requirements of neural networks.
Paper & contextCompressing and accelerating neural networks for deployment.
Han studies how to make neural networks efficient enough to deploy. These papers connect pruning and quantization with memory-efficient language-model inference and efficient generative models, considering both algorithms and the hardware executing them.
10 papers
Deep Compression combines pruning, weight quantization, and coding to reduce the storage requirements of neural networks.
Paper & contextAWQ uses activation information to choose weight quantization strategies that preserve model performance under compression.
Paper & contextAttention Sinks studies why a few initial tokens can preserve streaming language-model performance over long sequences, enabling efficient generation with limited memory.
Paper & contextSmoothQuant makes weight-and-activation quantization easier by redistributing the effect of activation outliers.
Paper & contextOnce-for-All trains a network that contains many usable subnetworks for different deployment constraints.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources