QLoRA: Efficient Finetuning of Quantized LLMs
QLoRA investigates fine-tuning quantized models with low-rank adapters to reduce memory requirements.
Paper & contextReducing the memory cost of large-model training and inference.
Dettmers studies how to reduce the memory and computing requirements of large models. The selection includes 8-bit optimization, LLM.int8(), and QLoRA, alongside research on practical deployment, reproducibility, and resource costs.
10 papers
QLoRA investigates fine-tuning quantized models with low-rank adapters to reduce memory requirements.
Paper & contextLLM.int8() reduces language-model memory requirements while treating unusually large feature values with higher precision.
Paper & context8-bit optimizers compress the running statistics used during training, reducing memory spent on optimizer state.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai, language models. Read the full paper, including the methods, experiments, and reported results.
Paper & contextSelected research in efficient ai. Read the full paper, including the methods, experiments, and reported results.
Paper & contextAn independent editorial profile. Inclusion does not imply Council membership or endorsement. Research is collaborative; coauthorship does not imply sole credit.
Selection & sources