Three prevalent themes in the discussion
-
Cost/value perception – Users repeatedly compare pricing, note that Opus 5.5‑High is half the cost of Opus 5‑High, and debate whether the “max” tier offers sensible bang‑for‑buck or is prohibitively expensive.
“Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.” – hglaser
“Opus 5.5's [Medium] actually can make sense dollar per intelligence/dollar per task wise… [Max]'s cost is completely unhinged.” – Someone1234 -
Overthinking and token limits of the “max” reasoning mode – Several commenters report that the maximum‑reasoning setting quickly exhausts its 128 k token budget on simple prompts, making it unsuitable for everyday work.
“I've failed twice to get 'Generate an SVG of a pelican riding a bicycle' to work with max, because in both cases it ran out of the 128,000 token budget while it was still reasoning about the problem.” – simonw
“'Max' is for benchmaxxing the intelligence metric and is not meant for use in productive work. Like drawing pelicans.” – samuelknight -
Skepticism about benchmark reliability and possible manipulation – Many doubt that the published scores reflect real‑world usefulness, citing performance drift after release and concern that indices are being gamed for marketing.
“I am begging you … to please stop posting this cringe… this index has Opus 5, one of the worst releases of 26, ahead of Astra. What information are we supposed to deduce from number having gone up?” – qsort
“Luna turned into drivel in essentially the same complexity of task… feels like it's being manipulated.” – seabass-salmon