Theme 1: Allegations that Chinese labs distilled US model reasoning traces
- “they are the authors of the well-known exploit to recover readable CoT from OpenAI and Anthropic models. They use that to find hints of distillation, by running a benchmark with a SotA model, recovering the CoT, then taking the first 1% of the CoT and running the open-source model as if that was the start of its own CoT.” — wongarsu
Theme 2: Skepticism about the evidence and alternative explanations
- “Did it occur to anyone that maybe the two sets of models were trained directly on the same solutions to the researchers' benchmark?” — nzeid
Theme 3: Ethical debate over copying, intellectual property, and openness
- “News flash: people who scraped the Internet without permission to build their product complain when something vaguely similar is done to them.” — CamperBob2