Open-Source LLMs You Can Host
How to choose a hostable open-weight model based on task fit, hardware limits, governance needs, and support burden rather than hype.
How to choose a hostable open-weight model based on task fit, hardware limits, governance needs, and support burden rather than hype.
What a private LLM deployment means in practice, when it makes sense, and how to compare managed private inference, self-hosting, and hybrid architectures.
TL;DR for operators A mixture-of-experts model can be too large for an inference budget even though only a small subset of its experts executes for each token. The inactive experts still have to reside somewhere. In the measured Mixtral deployment studied by Janati et al., removing four of eight experts per layer cut 4-bit memory from 24.2 GB to 12.3 GB and per-token latency from 40.3 ms to 25.2 ms on a single A100.1 ...
TL;DR for operators A production team may want every LoRA adapter to fit a small, uniform serving footprint. The usual response is to choose that small rank before training and optimize inside the resulting constraint. This paper shows that the training capacity and the deployment capacity do not always need to be identical. ...