1. Opus 5’s Slop & Modest Improvement
"reviews are overly pedantic, and that leads to feedback loops where each fix generates more feedback" — patwolf
"opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing" — Johnny_Bonk
"I almost prefer Opus 4.8. Opus 5.0 has the overly scholastic tone of Fable but without the intelligence." — roncesvalles
2. Need for Robust Long‑Term Code‑Quality Benchmarks
"maintainable is probably some high‑dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is" — dhorthy
"I'm not ready to blame Opus 5 for being stupid. Perhaps we have a prompt buried somewhere that's essentially asking it to be pedantic, and it's just obeying the prompt." — latentsea
"Really hoping that all of the attention you're bringing to the longitudinal sloppification of codebases makes it back to the labs..." — akurilin
3. Practical Harness & Cost Strategies
"the next thing on my radar is to try with a more realistic “factory‑shaped” harness where you have feedback from linters and other models after each coding episode" — dhorthy
"At least for Claude Code, putting “run /simplify at the end” in an “implement the plan” skill helps a little." — igregoryca
"Medium or low supposedly prevents Opus 5 from overthinking" — ValentineC