Theme 1: Cost and Value Considerations
- andreandre: “5 years of (200/month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm”
- simonw: “Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.”
Theme 2: Performance and Speed Benchmarks
- simonw (speed table):
Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
- beastman82: “The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.”
Theme 3: Hardware Architecture and Model Suitability
- tcdent: “A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform.”
- nacs: “People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.”
Theme 4: Ecosystem, Software, and Usability
- throwaway27448: “A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.”
- kokonokko1337: “Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously.”