Open-Source LLMs You Can Host

How to choose a hostable open-weight model based on task fit, hardware limits, governance needs, and support burden rather than hype.

March 16, 2026 · 7 min · Michelle

Deploy Your Own Private LLM

What a private LLM deployment means in practice, when it makes sense, and how to compare managed private inference, self-hosting, and hybrid architectures.

March 16, 2026 · 6 min · Michelle
Cover image

Probe Before You Prune: Measuring Which MoE Experts a Workload Can Lose

TL;DR for operators A mixture-of-experts model can be too large for an inference budget even though only a small subset of its experts executes for each token. The inactive experts still have to reside somewhere. In the measured Mixtral deployment studied by Janati et al., removing four of eight experts per layer cut 4-bit memory from 24.2 GB to 12.3 GB and per-token latency from 40.3 ms to 25.2 ms on a single A100.1 ...

September 6, 2026 · 9 min · Zelina
Cover image

Train Wide, Deploy Narrow: LoRA Rank Does Not Have to Be One Decision

TL;DR for operators A production team may want every LoRA adapter to fit a small, uniform serving footprint. The usual response is to choose that small rank before training and optimize inside the resulting constraint. This paper shows that the training capacity and the deployment capacity do not always need to be identical. ...

September 5, 2026 · 7 min · Zelina