Quantization
cs.LG · 2 papers
cs.LG · 2023
QLoRA: Efficient Finetuning of Quantized LLMs
VRAM requirement reduction: from >780GB to <48GBTim Dettmers, Artidoro Pagnoni +2
We present QLoRA, an efficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning…
Code· 11k★
cs.CL · 2024
The Era of 1-bit LLMs: BitNet b1.58
Energy Savings: 71.4xShuming Ma, Hongyu Wang +8
Recent research, such as BitNet, is paving the way for a new era of 1-bit Large Language Models (LLMs). In this work, we introduce a 1-bit LLM variant, namely BitNet b1.58, in…
Code· 22k★