🐺🐦‍⬛ LLM Format Comparison/Benchmark: 70B GGUF vs. EXL2 (and AWQ)

WolframRavenwolf@alien.top · 3 years ago

🐺🐦‍⬛ LLM Format Comparison/Benchmark: 70B GGUF vs. EXL2 (and AWQ)

ChiefBigFeather@alien.top · 3 years ago

This is difficult to evaluate. It could be that exl2 just breaks the translation layer.

ambient_temp_xeno@alien.top · 3 years ago

EXL2 5.0bpw was surprisingly doing much worse than GGUF Q2_K

Risitas.mov https://www.youtube.com/watch?v=QT13kk8HDDo

CosmosisQ@alien.top · 3 years ago

Hell yeah! Two days in a row! We need more people doing format comparisons and benchmarks in general. Again, thank you for all of your hard work, and keep 'em coming!

How would you say EXL2 subjectively compares to GGUF? Have you had the chance to roleplay with both formats outside of Voxta+VaM (i.e., in SillyTavern)? I ask because I’m sure the increased generation speed is more important than anything when using Voxta+VaM so it might be easier to compare their output quality in SillyTavern.

On that note, would you say you now prefer using lzlv (70B, EXL2) over OpenChat 3.5 (7B, GGUF) with Voxta+VaM?

tgredditfc@alien.top · 3 years ago

I have 2 gpus and AWQ never works for me on Oobabooga, no matter how I split the vRAM, oom in most of the cases.

thereisonlythedance@alien.top · 3 years ago

I had to split it something strange like 12/24GB to make it work. Even then I couldn’t get past 3K context.

panchovix@alien.top · 3 years ago

The major reason I use exl2 is speed, like on 2x4090 I get 15-20 t/s at 70b depending of the size, but GGUF I get like tops 4-5 t/s.

When using 3 gpus (2x4090+1x3090), it is 11-12 t/s at 6.55bpw vs GGUF Q6_K that runs at 2-3 t/s.

Though I agree with you, for model comparisons and such you need to have deterministic results and also the best quality.

If you can sometime, try 70b at 6bpw or more, IMO it is pretty consistent and doesn’t have issues like 5bpw/bits.

The performance hit is too much on multigpu systems when using GGUF. I guess if in the future the speed gets to the same level, I would use it most of the time.

a_beautiful_rhind@alien.top · 3 years ago

I’m surprised you get speeds so bad with GGUF. I get almost 9t/s on P40s and 18t/s on 3090.

GGUF is actually the fastest format until you load it up with context.

A couple of things have to be changed in cmakelists under vendor/llama.cpp if you’re using python

set(LLAMA_CUDA_MMV_Y        "2" CACHE STRING "llama: y block size for mmv CUDA kernels")
option(LLAMA_CUDA_FORCE_MMQ                  "llama: use mmq kernels instead of cuBLAS"         ON)

I have nvlink so this helps me. Since you don’t it still may help using direct communication via PCIE:

set(LLAMA_CUDA_PEER_MAX_BATCH_SIZE "8192" CACHE STRING

and since you’re using all new cards:

option(LLAMA_CUDA_F16                        "llama: use 16 bit floats for some calculations"   OFF)

Try out the FP16 support.

bullerwins@alien.top · 3 years ago

What motherboard do you have that can run 3x GPU’s?

candre23@alien.top · 3 years ago

GGUF I get like tops 4-5 t/s.

You’re doing something very wrong. I get better speeds than that on P40s with low context. Are you not using cublas?

easyllaama@alien.top · 3 years ago

‘The performance hit is too much on multigpu systems when using GGUF’

I agree. GGuF has multi GPU panelty. But it”s the most friendly to Apple silicons. I have same setup with you. one 4090 can run Xwin 13b at 40t/s. but when 2 cards present, it get only 1/4 of speed at 10t/s. So to get it fast, I have to flag CUDA device to single card while 2 cards present.

Since GGUF liks single GPU, those who have 3090/4090 will find 34B the best spot with the format.

Aaaaaaaaaeeeee@alien.top · 3 years ago

on 2.Xbpw, untick “add bos_token” avoiding the “cord string builder” looping

permalip@alien.top · 3 years ago

FYI, AWQ released 0.1.7 that fixes multi-GPU. Should alleviate OOM issues on multi-GPU, which became broken with newer versions of Huggingface libraries.

https://github.com/casper-hansen/AutoAWQ/releases/tag/v0.1.7

WolframRavenwolf@alien.top · 3 years ago

Oh, great news, once that’s in ooba, I’ll give it another try.

nsfw_throwitaway69@alien.top · 3 years ago

I wasn’t aware that Exl2 had issues with quality. Your tests seem to suggest that equivalent bpw in Exl2 produce worse results than in GGUF. I wonder why that is.

DataPhreak@alien.top · 3 years ago

The speeds don’t really surprise me. They’re going to take longer to load, but the math is about the same once they’re stood up.

w4ldfee@alien.top · 3 years ago

i run lzlv 2.4bpw without problems. make sure to disable bos token, then it should work way better.

Worldly-Mistake-8147@alien.top · 3 years ago

I’m probably going to ask something extremely basic, but why GPTQ isn’t an option? With OP’s double GPU he can run 4bit 32g with 8k context, and I was under impression that the quality loss is barely noticeable. Though I noticed it absolutely messes up numbers (math, or historical dates).

Ycros@alien.top · 3 years ago

It may be interesting to anyone running models across 2 3090s that in llama.cpp/koboldcpp there’s a performance increase if your two GPUs support peering with one another (check with nvidia-smi topo -p2p r) - it wasn’t working with my particular motherboard, so I installed an nvlink bridge and got a performance bump in token generation (an extra 10-20% with 70b, more with smaller models, except smaller models go much faster if you can fit them on one gpu).

I have no idea what the performance diff is between having a bridge and peering via pci-e if your system supports it. I also tested exl2 and there was no difference as I don’t think it implements any sort of peering optimisations.

lone_striker@alien.top · 3 years ago

For the 2.4bpw and 2.6bpw exl2 models, you have to change a setting in ooba to get them to generate coherent text. Disable this setting:

Add the bos_token to the beginning of prompts

https://preview.redd.it/4v8m7ciu0y1c1.png?width=356&format=png&auto=webp&s=785837b8466a3bcda3e49477424b7c377a8d542f

The very low bpw models need the above setting as well as being more strict with the prompt format. The higher bpw models are more flexible and can deal with prompt formats they were not specifically tuned for.

I would also set the VRAM for 2.4 to use only a single GPU. Spreading them out over two GPUs is not needed and will slow them down. That’s the main reason I generate 2.4 (and 2.6bpw) versions is to allow people with only a single 3090 or 4090 to run 70B models at full speeds. Though obviously quality will be lower than the higher-bit models. For 2.6bpw to fit on a single 24 GB VRAM GPU, you will need to enable the cache_8bit option.

WolframRavenwolf@alien.top · 3 years ago

Does 8-bit cache reduce quality or speed or what’s the disadvantage of it? (If it had none, it would be default, I assume.)

llama_in_sunglasses@alien.top · 3 years ago

GGUF k-quants are really good at making sure the most important parts of the model are not x bit but q6_k if possible. GPTQ and AWQ models can fall apart and give total bullshit at 3 bits while the same model in q2_k / q3_ks with around 3 bits usually outputs sentences.

ReturningTarzan@alien.top · 3 years ago

When you’re using non-instruct models for instruct-type questions, prompting is everything. For comparison, here are the first three questions put to Mistral-7B-instruct with correct prompt format at various bitrates up to FP16.

kpodkanowicz@alien.top · 3 years ago

Great work as always! Regarding Exl2 its sensitive to calibration dataset - probably the one that was used is not related to your tests. I.e. you can get higher scores in HumanEval even in 3 bits that you would get in transformers 8bit. I hope that this standard will get more popular and finetuners will do their own measurement file/quants using their dataset. Never seen q2 gguf doing better than exl2 unless i mixed rope config.

Edit - for anything higher than 4.25bit i usually use 8bit head

Model	Format	Quant	Offloaded Layers	VRAM Used	Primary Score	Secondary Score	Speed +mmq	Speed -mmq
lizpreciatior/lzlv_70B.gguf	GGUF	Q4_K_M	83/83	39362.61 MB	18/18	4+3+4+6 = 17/18
lizpreciatior/lzlv_70B.gguf	GGUF	Q5_K_M	70/83 !	40230.62 MB	18/18	4+3+4+6 = 17/18
TheBloke/lzlv_70B-GGUF	GGUF	Q2_K	83/83	27840.11 MB	18/18	4+3+4+6 = 17/18	4.20T/s	4.01T/s
TheBloke/lzlv_70B-GGUF	GGUF	Q3_K_M	83/83	31541.11 MB	18/18	4+3+4+6 = 17/18	4.41T/s	3.96T/s
TheBloke/lzlv_70B-GGUF	GGUF	Q4_0	83/83	36930.11 MB	18/18	4+3+4+6 = 17/18	4.61T/s	3.94T/s
TheBloke/lzlv_70B-GGUF	GGUF	Q4_K_M	83/83	39362.61 MB	18/18	4+3+4+6 = 17/18	4.73T/s !!	4.11T/s
TheBloke/lzlv_70B-GGUF	GGUF	Q5_K_M	70/83 !	40230.62 MB	18/18	4+3+4+6 = 17/18	1.51T/s	1.46T/s
TheBloke/lzlv_70B-GGUF	GGUF	Q5_K_M	80/83	46117.50 MB	OutOfMemory
TheBloke/lzlv_70B-GGUF	GGUF	Q5_K_M	83/83	46322.61 MB	OutOfMemory
LoneStriker/lzlv_70b_fp16_hf-2.4bpw-h6-exl2	EXL2	2.4bpw		11,11 -> 22 GB	BROKEN
LoneStriker/lzlv_70b_fp16_hf-2.6bpw-h6-exl2	EXL2	2.6bpw		12,11 -> 23 GB	FAIL
LoneStriker/lzlv_70b_fp16_hf-3.0bpw-h6-exl2	EXL2	3.0bpw		14,13 -> 27 GB	18/18	4+2+2+6 = 14/18
LoneStriker/lzlv_70b_fp16_hf-4.0bpw-h6-exl2	EXL2	4.0bpw		18,17 -> 35 GB	18/18	4+3+2+6 = 15/18
LoneStriker/lzlv_70b_fp16_hf-4.65bpw-h6-exl2	EXL2	4.65bpw		20,20 -> 40 GB	18/18	4+3+2+6 = 15/18
LoneStriker/lzlv_70b_fp16_hf-5.0bpw-h6-exl2	EXL2	5.0bpw		22,21 -> 43 GB	18/18	4+3+2+6 = 15/18
LoneStriker/lzlv_70b_fp16_hf-6.0bpw-h6-exl2	EXL2	6.0bpw		> 48 GB	TOO BIG
TheBloke/lzlv_70B-AWQ	AWQ	4-bit			OutOfMemory

🐺🐦‍⬛ LLM Format Comparison/Benchmark: 70B GGUF vs. EXL2 (and AWQ)

🐺🐦‍⬛ LLM Format Comparison/Benchmark: 70B GGUF vs. EXL2 (and AWQ)

My AI Workstation:

Observations:

Conclusion: