minus-squarepablines@alien.topBtoLocalLLaMA•What kind of specs to run local llm and serve to say up to 20-50 userslinkfedilinkEnglisharrow-up1·3 years agoHugging face text inference can handle concurrency you just need to power with gpus linkfedilink
minus-squarepablines@alien.topBtoLocalLLaMA•Rocket 🦝 - smol model that overcomes models much larger in sizelinkfedilinkarrow-up1·3 years agoWoooooooow! linkfedilink
Hugging face text inference can handle concurrency you just need to power with gpus