Top Themes in the Discussion
1. Alternative positional encoding & attention design
"Curious to see if it holds up at frontier scale." — gokohl
"KDA ... isn't really attention at all in any conventional sense." — thunderbird120
2. Scale and parameter‑count implications
"The number of active parameters is vastly different. Deepseek CEO hinted that he estimates it as an order of magnitude difference in one of his recent interviews." — porridgeraisin
3. Inference‑level caching trade‑offs
"It results in up to 1023 additional input (cache miss) tokens per inference." — samuelknight