← Flash Papers
QLoRA: Efficient Finetuning of Quantized LLMs
Share
Share this Flash Paper
QLoRA: Efficient Finetuning of Quantized LLMs
https://flashpapers.ai/p/ntq7uuag9g
Copy link
Post to X
Post to Bluesky
Share on LinkedIn
Submit to Hacker News
Submit to Reddit
Email a link
Source PDF ↗
Open standalone ↗
Make your own
Make your own paper
Tim Dettmers, Artidoro Pagnoni +2 more
cs.LG
2023
arXiv:2305.14314
Code
· 11k★
· updated 2y ago
Details ▸
Details ▴
VRAM requirement reduction:
from >780GB to <48GB
Guanaco 65B relative performance:
99.3%
· Vicuna Benchmark
Double Quantization savings:
0.37 bits/parameter
VRAM requirement reduction:
from >780GB to <48GB
Guanaco 65B relative performance:
99.3%
· Vicuna Benchmark
Double Quantization savings:
0.37 bits/parameter
Concepts:
LoRA