The title, pretty much.
I’m wondering whether a 70b model quantized to 4bit would perform better than a 7b/13b/34b model at fp16. Would be great to get some insights from the community.
The title, pretty much.
I’m wondering whether a 70b model quantized to 4bit would perform better than a 7b/13b/34b model at fp16. Would be great to get some insights from the community.
about 44 GB
Using Q3, you can fit it in 36GB (I have a weird combo of RTX 3060 with 12GB and P40 with 24GB and I can run a 70B at 3bit fully on GPU).
So you have 2 GPUs on single m/b… and the llama.cpp thing knows to use both? Does this work with AMD GPUs too?
Yes llama.cpp will automatically split the model to work across GPUs. You can also specify how much of the full model should be on each GPU.
Not sure on AMD support but for nvidia it’s pretty easy to do.
44GB of GPU VRAM? WTH GPU has 44GB other than stupid expensive ones? Are average folks running $25K GPUS at home? Or those running these like working for company’s with lots of money and building small GPU servers to run these?
Dual 3090/4090s. Still pricey as hell, but not out of reach for some folks.
So anyone wanting to play around with this at home, has to expect to drop about 4K or so for GPUs and a setup?
I can get 2 3090 for 1200€ here on the second-hand market