1️⃣ Qwen 3.8‑27B rivals much larger frontier models
“Same score as the latest DeepSeek Flash 0731 which has 284B parameters! (13B active)” – bertili
“It’s basically a match for larger frontier models — hundreds of times larger — in the narrow domain of mathematical and logical reasoning problems.” – CamperBob2
“3.8 actually performs slightly worse than 3.6 on AA‑Omniscience Accuracy, implying a trade‑off in world knowledge for other capabilities.” – anana_
2️⃣ Token‑efficiency & deployment trade‑offs
“It also produces nearly twice as many tokens per task as 3.6, which may be required to achieve correctness at this parameter size.” – anana_
“The token usage is 2.3× GPT‑Luna Max and almost 2× Kimi K3!” – phsource
“Qwen models are slower in tokens/s and use more tokens per task, which can hurt local deployment.” – petu
3️⃣ Benchmarks, “bench‑maxxing” & real‑world relevance
“The biggest untapped market is pure agentic models built for tool‑calling and non‑hallucination instead of memorizing facts.” – tancop
“We don’t actually know how large they are, actually.” – johnnyApplePRNG
“This model is nowhere close to other models in that score range.” – bertili (highlighting skepticism about AA rankings).