Right now it seems we are once again on the cusp of another round of LLM size upgrades. It appears to me that having 24gb VRAM gets you access to a lot of really great models, but 48gb VRAM really opens the door towards the impressive 70B models and allows you to nicely run the 30B models. However, im seeing more and more 100B+ models being created that push the 48 gb VRAM specs down into lower quants if they are able to run the model at all.

this is in my opinion is big, because 48gb is currently the magic number for in my opinion consumer level cards, 2x 3090’s or 2x 4090s. adding an extra 24gb to a build via consumer GPUs turns into a monumental task due to either space in the tower or capabilities of the hardware AND it would put you at 72gb VRAM putting you at the very edge of the recommended VRAM for the 120GB 4KM models.

I genuinely don’t know what i am talking about and i am just rambling, because i am trying to wrap my head around HOW to upgrade my vram to load the larger models without buying a massively overpriced workstation card. should i stuff 4 3090’s into a large tower? settle up 3 4090’s in a rig?

how can the average hobbyist make the jump from 48gb to 72gb+?

is taking the wait and see approach towards nvidia dropping new scalper priced high VRAM cards feasible? Hope and pray for some kind of technical magic that drops the required VRAM while simultaneously keeping quality?

the reason i am stressing about this and asking for advice is because the quality difference between smaller models and 70B models is astronomical. and the difference between the 70B models and the 100+B models is a HUGE jump too. from my testing it seems that the 100B+ models really turn the “humanization” of the LLM up to the next level, leaving the 70B models to sound like…well… AI.

I am very curious to see where this gets to by the end of 2024, but for sure… i won’t be seeing it on a 48gb VRAM set up.

  • DominicanGreg@alien.topOPB
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 year ago

    Parts wise, a threadripper + ASUS Pro WS WRX80E-SAGE SE WiFi II is already a 2k price floor.

    each 4090 is 2-2.3k

    each 3090 is 1-1.5k

    so building a machine from scratch will run you easily 8-10k off 4090’s and 6-8k off 3090’s. If you already have some GPUS or parts you would still problaly need 2 or more extra gpu’s plus the space and power to run them.

    to my specific situation i would have to grab the treadripper, mobo, a case, ram, 2 more cards im looking at potentially 5-7k worth of damage. OR… pay 8.6k for a mac pro m2 and get an entire extra machine to play with.

    There’s definitely an entire Mac Pro M3 series on the way considering they just released the laptops, it’s only a matter of time for them to shoot out the announcements. So i would definitely feel a bit peeved if i bought the M2 tower only for a month or two later apple to release the m3 versions.