Models Megathread #2 - What models are you currently using?

Technical_Leather949@alien.top · 2 years ago

Models Megathread #2 - What models are you currently using?

HvskyAI@alien.top · 2 years ago

I’m late to the party on this one.

I’ve been loving the 2.4BPW EXL2 quants from Lone Striker recently, specifically using Euryale 1.3 70B and LZLV 70B.

Even at the smaller quant, they’re very capable, and leagues ahead of smaller models in terms of comprehension and reasoning. Min-P sampling parameters have been a big step forward, as well.

The only downside I can see is the limitation to context length on a single 24GB VRAM card. Perhaps further testing of Nous-Capyabara 34B at 4.65BPW on EXL2 is in order.

FullOf_Bad_Ideas@alien.top · 2 years ago

Remember to try 8-bit cache If you haven’t yet, it should get you to 5.5k tokens context length.

You can get around 10-20k context length with 4bpw yi-34b 200k quants on single 24GB card.