abandonedexplorer@alien.topB to

LocalLLaMAEnglish · 2 years ago

Where and how to run Goliath 120b GGUF with good performance?

1

Where and how to run Goliath 120b GGUF with good performance?

abandonedexplorer@alien.topB to

LocalLLaMAEnglish · 2 years ago

I am talking about this particular model:

https://huggingface.co/TheBloke/goliath-120b-GGUF

I specifically use: goliath-120b.Q4_K_M.gguf

I can run it on runpod.io on this A100 instance with “humane” speed, but it is way too slow for creating long form text.

https://preview.redd.it/fz28iycv860c1.png?width=350&format=png&auto=webp&s=cd034b6fb6fe80f209f5e6d5278206fd714a1b10

These are my settings in text-generation-webui:

https://preview.redd.it/vw53pc33960c1.png?width=833&format=png&auto=webp&s=0fccbeac0994447cf7b7462f65d79f2e8f8f1969

Any advice? Thanks

Chat

panchovix@alien.topB
link
fedilink
English
arrow-up
1·
2 years ago
I tested 4K and it worked fine at 4.5bpw. Max will be prob about 6k. I didn’t use 8bit cache

Now 4.5bpw is kinda overkill, 4.12~ bpw is like 4bit 128g gptq, and that would let you use a lot more context.
- Dead_Internet_Theory@alien.topB
  link
  fedilink
  English
  arrow-up
  1·
  2 years ago
  That is awesome. What kind of platform do you use for that 3 GPUs setup?