1. Benchmark credibility and manipulation
Commenters argue that the Artificial Analysis scores are arbitrary—weights are shifted to favor certain models, and the numbers look like they come from a template.
- “The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that.” — seahorseemoji
- “The main AA benchmark keeps changing, and had to be radically changed when Astra came out …” — SyneRyder
- “Yes, it seems they have a template that they fill with numbers.” — theanonymousone
2. Usage limits, pricing, and perceived intelligence decline
Many users report that commercial plans (e.g., OpenAI/Astra) have tighter limits, slower speeds, and a noticeable drop in model usefulness, pushing them toward cheaper alternatives.
- “OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining…” — Gareth321
- “I’ve stopped using Astra entirely … I get roughly 1~2 days [of work] out of a weekly $200 plan.” — unsupp0rted
- “On the $100 plan I can burn my entire week’s budget in a few hours easily with Astra. It's borderline unusable.” — sauwan
3. Cost‑effectiveness of newer models (especially Chinese ones)
Discussion highlights that models like MiMo‑V2.6 and DeepSeek are far cheaper to train and run while delivering comparable performance, making their pricing a key advantage.
- “Per Xiaomi, MiMo v2.6 training run cost $3.47 m. A far cry from the estimated costs ($100m+) for the Big 5…” — ignoramous
- “It is an impressive model … But it's pricing is where it really shines.” — dom96
- “I seriously doubt salaries are included. It must be just the electricity and GPU costs.” — f6v