1. Concerns about reproducibility and the ease of gaming benchmarks
Many commenters worry that the current benchmarks are opaque, non‑reproducible, and can be “benchmaxxed” or manipulated.
- “So TL;DR benchmarking in a completely non‑reproducible manner ?” – traceroute66
- “You're giving up transparency for it being harder to game.” – kadoban
2. A noticeable gap between benchmark scores and real‑world user experience
Users frequently report that a model’s benchmark ranking does not match how well it works for them in practice.
- “The fact that gemini 3.8 flash is so high up there just tells you this is an awful benchmark. Try and use gemini 3.8 yourself for any real world work and you'll see it's terrible.” – bdlowery
- “Hard disagree. I use 3.8 flash in Antigravity a lot, and thoroughly prefer it to most Pro‑class models… It solved a problem I couldn't solve for weeks in under 6 hours.” – tucnak
3. Benchmarks can still be useful for relative comparison when conditions are held constant
Despite the flaws, several participants argue that benchmarks retain value for comparing models under the same setup or for building trust over time.
- “In theory, as long as all the models are doing the same thing with the same tools, it's at least useful to see how they stack up against each other right now.” – deepwoods
- “Anecdotally, +1. I’d also say this benchmark matches my experiences and how much I trust the model output.” – retrobox