1. Benchmarking skepticism / “benchmaxing”
Many commenters doubt the advertised scores, pointing out the steep drop from older benchmarks to the newer Terminal‑Bench 4 as evidence of over‑fitting.
- “Postalcoder: … this is the definition of bench maxxing.”
- “enraged_camel: Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.”
- “nullbio: Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity.”
2. Model origins, distillation, and openness concerns
A recurring thread is that SWE‑2 appears to be a post‑trained/distilled version of the Chinese Kimi K3 model, raising questions about novelty, weight availability, and whether it’s just a re‑package.
- “Tsarp: SWE-2 is post-trained from Kimi K3.”
- “airstrafer: I guess I don't. Does post‑training from another (larger) model not fall under the umbrella of distillation? I’d imagine it leads to the same spiky‑ness issues…?”
- “FergusArgyll: Distilling you don't have the actual model weights of the teacher… Fine tuning you have the actual model weights of the original model…”
- “xlbuttplug2: I presume post training is significantly easier than the distillation/training the top Chinese labs are doing…”
3. Platform lock‑in and usability concerns (Devin‑centric access)
Several users criticize the requirement to use Cognition’s own CLI/Devin ecosystem, saying it creates friction and limits easy evaluation alongside other models.
- “scronkfinkle: Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent…”
- “wren6991: Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are not an interchangeable commodity already?”
- “randomblock1: I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage…”
- “anthonypasq: they are an agent company not a model provider, is this that difficult to comprehend?”