Three dominant themes in the discussion
| Theme | Supporting quote |
|---|---|
| 1. Real‑world performance limits – Many users stress that token‑throughput numbers are misleading because the bottleneck is often prefill or memory swapping, and current speeds on consumer hardware feel “not great.” | “this decode times don't really tell the whole story, because prefill becomes the bottleneck” — brrrrrm |
| 2. Incremental progress is expected – The community believes that each small speed gain (e.g., 3 t/s → 6 t/s) will stack up toward much higher rates, indicating ongoing progress despite current constraints. | “This is how progress happens, someone gets to 3t/s, the next person gets to6/s and eventually we get to 100t/s” — fsuts |
| 3. Economics nudging inference to the cloud – Several commenters argue that the cost of serving many users makes centralized servers the more sensible model, and local LLMs will likely stay a niche or “background” use‑case. | “I suspect the economics favor centralized servers, if you only look at the aggregated cost to serve X number of users' tokens” — anon373839 |
These three themes capture the prevailing sentiment: a realistic view of today’s speed limits, confidence that improvements will continue, and an expectation that most heavy‑weight LLM work will remain server‑based for the foreseeable future.