1. Extreme single‑core optimization on Zen 3
"I recently took on the challenge of squeezing every drop of performance out of a single AMD Zen 3 core, successfully reaching a blazing‑fast 85.3 GFLOPS in FP32 matrix multiplication." – houslast
2. CPU vs GPU performance & the memory‑bandwidth bottleneck
"For comparison, the best performing GPUs today can do FP32 at > 100 TFLOP/s." – ranger_danger
"The real limit is almost certainly memory bandwidth, not flops." – pixelpoet
3. Future architectures, cost‑performance outlook & analog alternatives
"Could alternative CPU architectures like ternary/quaternary and/or analog etc be better at matrix multiplication?" – Razengan
"Analog is certainly better, so long as you're OK with some noise." – entropicdrifter