1. Asynchronous/local learning reduces global coordination but can trade off efficiency
- “The win with asynchronous techniques like NPC is that you do not need the extreme co‑ordination that backprop requires and hence should be computationally much easier given the right device.” – vatsachak
- “Much easier given the right device” … “the price of ‘not having backprop’ is usually expending more FLOPs, getting worse sample efficiency, etc.” – ACCount39
2. Coordination cost depends heavily on the substrate (GPU vs. brain)
- “Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent … coordinating them can get less natural … When wiring is expensive, coordination can become expensive in turn.” – ACCount39
- “I mean co‑ordination requires energy though. The brain wattage looks at GPUs and says ‘skill issue’.” – vatsachak
3. Practical deployment faces challenges in distributed training, continual learning, and catastrophic forgetting; hybrid/modular ideas may help
- “Imagine we did that, split up a model layers as A->B->C… This is in contrast to mining bitcoins … which doesn’t require any coordination from miners…” – janalsncm
- “DUST does have an advantage … because it doesn’t have to save a ton of intermediate state … There are many other issues that this algorithm does not address … catastrophic forgetting.” – nbutton762
- “Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint … unlock further gains?” – polyomino