3 Dominant Themes from the HN Thread
| # | Theme | Supporting Quotations |
|---|---|---|
| 1 | Kimi K3’s novel architecture (KDA vs. MLA) | • “I believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA)… previous large Kimi models had only MLA layers.” – GaggiX • “It’s not the same KDA as used in Kimi Linear, though.” – yorwba • “The main contribution of the K3 paper is Stable LatentMoE. … more balanced expert selection strategy.” – throwa356262 (linked to https://arxiv.org/abs/2607.24653) |
| 2 | Ethics & IP around model distillation | • “Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans.” – fwip • “This is a nakedly hypocritical stance, but completely understandable from a company‑needs‑to‑make‑money standpoint.” – fwip • “Google is offering distillation as a paid product.” – verdverm (link to docs) |
| 3 | Scaling laws & emergent capability debates | • “Is it weird that a 1 trillion‑parameter model suddenly can conjure up counter‑examples for the Jacobian conjecture while a 1 M‑parameter model can’t?” – pooyamo • “The Bitter Lesson shows that general‑purpose methods keep scaling with compute, even if we don’t understand the underlying algorithm.” – pachev (refers to https://en.wikipedia.org/wiki/Bitter_lesson) • “Training a frontier model is not just brute‑force; gradient descent compresses trillions of tokens without explicit brute force.” – mohsen1 |
All quotations are reproduced verbatim with double quotes and proper HTML‑entity fixes.