Theme 1: CPU vs. GPU efficiency for matrix‑multiplication workloads
- “I wonder what would 'compute' mean if cpus were more efficient at matrix multiplication …” – tolugenius
- “… vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.” – actionfromafar
- “CPUs are designed to perform a handful of operations at one time… GPUs are designed to process a large number of calculations at once …” – rhdunn
Theme 2: Memory (VRAM/RAM) limits model size and batch size
- “The other issue when training models … is the amount of VRAM (or RAM for CPUs) available… This affects things like batch size and the size of model that can be trained or fine‑tuned.” – rhdunn
- “A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training …” – rhdunn
Theme 3: Hardware/software feedback loop and the need for more flexible, general‑purpose architectures
- “As the video points out, there is a hardware/software feedback loop at play here… I predict that a set of general purpose CPUs – non shared memory at this scale – would be a much better use of transistor/power.” – librasteve
- “… you can code that at high level give a CSP style approach such as https://bil-lang.org” – librasteve