Three Prevalent Themes in the Discussion
1. Cost‑and‑Performance Claims
Many commenters focus on the asserted token savings and accuracy gains of the Strands harness versus other setups.
- “Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42. … only the harness is different with a ~5 point difference in accuracy while costing significantly less.” – CharlieDigital
- “With Fable 5, Strands harness cost 77% less than Claude Code and scored higher on Terminal Bench 2.1.” – seizethecheese
2. Debate Over Excluding Vanilla Pi
A recurring point is whether it is fair to compare Strands (or Oh‑My‑Pi/Deepseek) without including the lightweight base Pi harness.
- “They have oh‑my‑pi in the benchmark but not pi, but those are very different animals. pi is lightweight out of the box so has very little start‑time overhead.” – jsw97
- “Base Pi doesn’t seem like an apt comparison here. Strands comes with MCP servers, subagents, and web fetch tools built in. To get those on Pi, you have to add addons …” – cobolcomesback
- “Why is Pi not in the benchmarks? Deepseek beats Strands and its built on Pi …” – theturtletalks
3. Skepticism About Marketing‑Driven Benchmarks
Several users criticize the hype, calling the results “voodoo” or pointing out that benchmarks are saturated and prone to selective interpretation.
- “This is at least the fourth time I’ve seen a project hit front page with a ‘save money with same score on saturated benchmark’ claim.” – seizethecheese
- “Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches …” – avaer
- “I get what you're saying, but their graphic … shows a significant difference …” – CharlieDigital (highlighting the tension between skepticism and the presented data)