Project ideas from Hacker News discussions.

What if we stopped using GPUs? [video]

📝 Discussion Summary (Click to expand)

Theme 1: CPU vs. GPU efficiency for matrix‑multiplication workloads
- “I wonder what would 'compute' mean if cpus were more efficient at matrix multiplication …” – tolugenius
- “… vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.” – actionfromafar
- “CPUs are designed to perform a handful of operations at one time… GPUs are designed to process a large number of calculations at once …” – rhdunn

Theme 2: Memory (VRAM/RAM) limits model size and batch size
- “The other issue when training models … is the amount of VRAM (or RAM for CPUs) available… This affects things like batch size and the size of model that can be trained or fine‑tuned.” – rhdunn
- “A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training …” – rhdunn

Theme 3: Hardware/software feedback loop and the need for more flexible, general‑purpose architectures
- “As the video points out, there is a hardware/software feedback loop at play here… I predict that a set of general purpose CPUs – non shared memory at this scale – would be a much better use of transistor/power.” – librasteve
- “… you can code that at high level give a CSP style approach such as https://bil-lang.org” – librasteve


🚀 Project Ideas

CPU-ML Matrix Kernel Tuner

Summary

  • Auto‑tunes GEMM and related matrix multiplication kernels for modern CPUs (AVX‑512, AMX, SME) to maximize throughput for LLM workloads.
  • Generates specialized kernels that can be dropped into PyTorch/TensorFlow via a simple import, reducing reliance on GPU‑only ops.

Details

Key Value
Target Audience ML researchers and engineers who want to train/fine‑tune LLMs on CPU‑only machines with large RAM
Core Feature Instruction‑set‑aware kernel generator with empirical search, caching, and fallback to BLAS
Tech Stack C++, LLVM/MLIR, Python bindings, Google Benchmark, optional CUDA‑free backend
Difficulty Medium
Monetization Revenue-ready: SaaS subscription for premium kernel profiles and support

Notes

  • HN users lamented the lack of efficient CPU matrix multipliers (rh­dunn: “Having dedicated matrix multiplication instructions would still lock that to how many matrix‑capable ALUs there are”). This tool directly addresses that by squeezing more performance out of existing ALUs.
  • Could spark discussion on open‑source kernel sharing and enable CPU‑centric ML benchmarks.

LLM‑CPU Memory Planner

Summary

  • Predicts memory footprint (weights, gradients, activations, optimizer states) for training or fine‑tuning any transformer model on CPU given model size, precision, batch size, and sequence length.
  • Recommends memory‑saving techniques (gradient checkpointing, LoRA/QLoRA, weight quantization, activation offloading) and estimates resulting speed/accuracy trade‑offs.

Details

Key Value
Target Audience Practitioners planning CPU‑based LLM experiments, especially those limited by RAM (as noted by rhdunn and unsloth.ai links)
Core Feature Interactive CLI/web UI that takes model config and hardware specs, outputs memory breakdown and optimization suggestions
Tech Stack Python, FastAPI (for web), PyTorch meta‑modules for shape analysis, React frontend
Difficulty Low
Monetization Hobby

Notes

  • Commenters asked “I don’t know what sized model you could train on 64GB/128GB RAM via a CPU.” This tool answers that question concretely.
  • Provides a practical utility for hobbyists and small labs, encouraging more CPU‑focused experimentation.

Distributed CPU Training Fabric

Summary

  • A lightweight framework that orchestrates model‑parallel training across many CPU nodes using MPI/RDMA, exploiting each node’s large RAM to shard model layers, optimizer states, and activation checkpoints.
  • Handles automatic gradient synchronization, pipelining, and fault tolerance, enabling training of models that exceed single‑node GPU memory limits.

Details

Key Value
Target Audience Research teams and companies with CPU clusters (or cloud VMs) seeking to train large LLMs without GPUs
Core Feature Distributed training API compatible with HuggingFace Transformers and PyTorch, with automatic stage partitioning and communication scheduling
Tech Stack C++/Rust core, MPICH or Open MPI, libfabric for RDMA, Python bindings, optional Kubernetes operator
Difficulty High
Monetization Revenue-ready: Enterprise license with support tiers

Notes

  • rhdunn highlighted scalability limits of CPUs; this project turns that limitation into a strength by leveraging many CPUs together, a point echoed by librasteve’s “general purpose CPUs – non shared memory at this scale”.
  • Could stimulate discussion on alternative hardware‑software co‑design for ML, mirroring the thread’s focus on rethinking compute.

Read Later