Project ideas from Hacker News discussions.

Dust: Pretraining Transformers Without Backpropagation

📝 Discussion Summary (Click to expand)

1. Asynchronous/local learning reduces global coordination but can trade off efficiency
- “The win with asynchronous techniques like NPC is that you do not need the extreme co‑ordination that backprop requires and hence should be computationally much easier given the right device.” – vatsachak
- “Much easier given the right device” … “the price of ‘not having backprop’ is usually expending more FLOPs, getting worse sample efficiency, etc.” – ACCount39

2. Coordination cost depends heavily on the substrate (GPU vs. brain)
- “Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent … coordinating them can get less natural … When wiring is expensive, coordination can become expensive in turn.” – ACCount39
- “I mean co‑ordination requires energy though. The brain wattage looks at GPUs and says ‘skill issue’.” – vatsachak

3. Practical deployment faces challenges in distributed training, continual learning, and catastrophic forgetting; hybrid/modular ideas may help
- “Imagine we did that, split up a model layers as A->B->C… This is in contrast to mining bitcoins … which doesn’t require any coordination from miners…” – janalsncm
- “DUST does have an advantage … because it doesn’t have to save a ton of intermediate state … There are many other issues that this algorithm does not address … catastrophic forgetting.” – nbutton762
- “Even though this is way more expensive than backprop, could a hybrid approach where you fine tune an existing checkpoint … unlock further gains?” – polyomino


🚀 Project Ideas

Generating project ideas…

AsyncTrain: PyTorch Extension for Local Learning Rules

Summary

  • Enables replacement of backprop with biologically‑inspired local update rules (e.g., Neural Predictive Coding) to reduce coordination overhead in distributed training.
  • Core value proposition: faster, more energy‑efficient training on existing GPU clusters by eliminating global gradient synchronization.

Details

Key Value
Target Audience ML researchers and engineers experimenting with alternative training algorithms
Core Feature Drop‑in modules that implement local, asynchronous weight updates compatible with standard nn.Module
Tech Stack PyTorch, CUDA extensions, Python
Difficulty Medium
Monetization Hobby

Notes

  • HN commenters highlighted the need for “asynchronous techniques like NPC” that avoid expensive coordination (vatsachak) and the desire for “online learning” that is cheap and stable (ACCount39).
  • Provides a practical testbed for the ideas discussed, potentially sparking new submissions to ML conferences and open‑source collaborations.

NPC‑Sim: Event‑Driven Simulator for Asynchronous Neural Networks

Summary

  • Cycle‑accurate, event‑based simulator for spiking and predictive‑coding networks to estimate energy, bandwidth, and coordination costs.
  • Core value proposition: lets hardware architects and algorithm designers quickly evaluate trade‑offs before building custom ASICs or FPGA boards.

Details

Key Value
Target Audience Hardware designers, neuromorphic engineers, ML systems researchers
Core Feature Configurable neuron models, asynchronous update semantics, power and traffic profiling
Tech Stack C++ (SystemC), Python bindings, optional GUI with Dear ImGui
Difficulty High
Monetization Hobby

Notes

  • Discussion noted that “co‑ordination requires energy” and that the brain’s low wattage looks at GPUs and says “skill issue” (vatsachak); a simulator would let users quantify those energy differences.
  • Enables concrete experiments to back up claims about NPC gradients converging to backprop within a regime, addressing skeptics who ask for proof of scaling to billions of parameters.

ChunkMoE: Decentralized Mixture‑of‑Experts Training Framework

Summary

  • Framework for training Mixture‑of‑Experts models where each expert chunk updates independently via gossip‑based parameter exchange, eliminating global synchronization steps.
  • Core value proposition: scales training to massive models with minimal network bandwidth and deterministic convergence properties.

Details

Key Value
Target Audience Large‑scale ML practitioners, teams training trillion‑parameter models
Core Feature Asynchronous expert training with periodic, low‑overhead state merging
Tech Stack Rust (for networking), PyTorch bindings, Apache Arrow for data exchange
Difficulty High
Monetization Revenue-ready: SaaS subscription for managed clusters + enterprise support

Notes

  • Commenters pointed out the pain of moving gigabytes of activations between layers in distributed backprop (janalsncm) and the advantage of widthwise parallelism in MOE (vatsachak); ChunkMoE directly addresses those bottlenecks.
  • Offers a practical path to test the hypothesis that “0th order methods will parallelize better than backprop especially along depth” (SerdarGl) while providing a clear monetization route via cloud‑hosted training services.

Read Later