“Some models (like DeepSeek’s) that are mixture-of-experts with many layers thus require large batch sizes and high latency, otherwise throughput drops off a cliff.”

Mixture-of-experts models only run cheap when thousands of requests are batched together to keep every expert busy. One user running one model locally gets none of that. That makes cheap inference a property of scale, not of the model. The open weights do not come with the economics.