Local generation not utilising GPU?

AlternativeParfait47@alien.top · 1 year ago

Local generation not utilising GPU?

cndvcndv@alien.top · 1 year ago

I am not sure what installing llama means. There are different ways of running llama. But if the program you installed is supposed to utilize gpu, it could be a cuda issue.

Lup0Grigi0@alien.top · 1 year ago

If you have installed or use Oogabooga tex-generation-webui download a model that has ben quantized for nVidia GPU. Those are the models with GTPQ and the newer AWQ suffixes.
On hugging face, the user “thebloke” has aggregated dozens and dozens, maybe hundreds, of models.
the youtube channel Aitrepreneur a couple good videos on installing ooga and how to run the GPU quantized models

FlishFlashman@alien.top · 1 year ago

What software are you using to run LLaMA and Stable Diffusion?

What version of the LLaMA model are you trying to run? How many parameters? What quantization?