Project ideas from Hacker News discussions.

OpenTPU – An open-source AI accelerator, developed by AI

📝 Discussion Summary (Click to expand)

Summary of prevalent themes

  • AI‑driven hardware design & recursive self‑improvement

    “After using AI to develop risc‑v CPU cores, the same technique was used for developing openTPU… The TPU started able to produce only a few tokens per second and through a recursive self‑improvement loop got to 80+ tok/sec on the smaller models.” – fsbonetto

  • Hardware feasibility challenges (FPGA limits, memory bandwidth, ASIC vs FPGA)

    “For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.” – fsbonetto
    “The current largest FPGA… has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75 M).” – LoganDark

  • Economic and obsolescence concerns (model turnover outpaces chip development)

    “Model SOTA moves faster than chips can be designed or produced. You'd need to commit to a particular model for years to get payoff while still burning buckets of money producing new SOTA models to keep up with the competition.” – zdragnar

  • Societal impact: job displacement and race‑to‑bottom pressures from cheap, fast inference

    “At this point, LLM's are 'good enough' for all kinds of tasks. Instead of making them more capable, now the efforts are making them smaller and cheaper. All aboard! We're racing to the bottom now.” – HoldOnAMinute


🚀 Project Ideas

FPGA-in-the-Loop Validation for LLM-Generated Hardware

Summary

  • Cloud service that takes RTL produced by LLMs (e.g., from natural‑language prompts) and runs cycle‑accurate FPGA emulation with benchmark inference of known LLMs to verify correctness before tapeout.
  • Provides confidence that AI‑designed hardware works in silicon, solving the “simulation not accurate w.r.t reality” problem.

Details

Key Value
Target Audience Hardware designers, AI researchers experimenting with LLM‑generated chip designs
Core Feature Automated validation pipeline: LLM prompt → RTL synthesis → FPGA bitstream → run Llama2/Qwen inference → performance/power report
Tech Stack Verilator, SymbiYosys, AWS F1 instances, Python orchestration, Docker
Difficulty Medium
Monetization Revenue-ready: pay‑per‑validation run ($X per hour of FPGA time)

Notes

  • Addresses sailingparrot’s concern: “Getting an LLM to design something in its own simulator that is not accurate w.r.t reality is not useful nor terribly impressive.”
  • Gives the OpenTPU community a quick way to check AI‑generated improvements before committing to fab.

Model-on-Chip Inference as a Service

Summary

  • Offers pre‑burned ASIC (or FPGA) inference engines containing popular open‑weight LLMs (e.g., Qwen, Gemma) via a low‑latency API.
  • Enables developers to access cheap, fast inference for “good enough” models without managing hardware themselves.

Details

Key Value
Target Audience Startups, indie developers, edge‑device makers seeking affordable LLM inference
Core Feature API endpoints for text generation on dedicated AI accelerator hardware with guaranteed sub‑10 ms/token latency for 8B models
Tech Stack Custom ASIC/FPGA bitstream, FastAPI/gRPC, Docker, Kubernetes for orchestration
Difficulty High (hardware development)
Monetization Revenue-ready: per‑token pricing (e.g., $0.0001 per 1k tokens) or monthly reserved capacity

Notes

  • Mirrors jjcm’s comment: “I would happily use an opus 4.7 at 15k tokens per second.” and Cerebras serving older models at high speed.
  • Solves the demand for cheap inference expressed by commenters who want to run “good enough” models at high throughput.

Responsible AI Hardware Design Sandbox

Summary

  • Development environment that logs every LLM prompt used to generate hardware, visualizes design changes, and enforces safety constraints (e.g., blocks crypto‑mining or weaponizable circuits) before allowing synthesis.
  • Creates an audit trail for accountability and reduces risk of unintended AI behavior.

Details

Key Value
Target Audience AI researchers and chip designers using LLMs for hardware generation
Core Feature Prompt tracking, design diff visualization, constraint checking (power limits, forbidden IP blocks), compliance audit trail
Tech Stack JupyterLab extension, LLM APIs, Verilog/SystemVerilog parsers, Open Policy Agent (OPA), Git integration
Difficulty Medium
Monetization Hobby (open core) or Revenue-ready: enterprise license for audit/compliance features

Notes

  • Responds to the responsibility debate sparked by pixl97 and dpoloncsak: “It does what it finds it needs to do to achieve the goal defined in the prompt.”
  • Provides the oversight that commenters felt was missing when LLMs act autonomously.

Compute-in-Memory (CIM) SDK for LLM Inference

Summary

  • High‑level SDK and compiler that maps LLM layers onto compute‑in‑memory architectures, abstracting memory bandwidth concerns and letting ML engineers target emerging CIM chips without writing HDL.
  • Includes quantization, microcode generation, and bandwidth‑aware simulation.

Details

Key Value
Target Audience ML engineers wanting to deploy LLMs on next‑gen AI accelerators
Core Feature Takes PyTorch/TensorFlow model → applies quantization → generates CIM‑compatible microcode → estimates bandwidth usage → offers functional simulation
Tech Stack MLIR, LLVM, Python, CIM simulator (e.g., NeuroSim, Gemini), PyTorch bindings
Difficulty High
Monetization Revenue-ready: subscription for cloud‑based CIM simulation access ($XX/mo) or perpetual compiler license

Notes

  • Directly addresses fsbonetto’s observation: “The bottleneck, for inference at least, is memory bandwidth.”
  • Enables the compute‑in‑memory approach hinted at by fnordpiglet and cestith, turning a hardware limitation into a software‑friendly flow.

OpenTPU Integration Kit

Summary

  • Open‑source hardware/software bundle that simplifies integrating the OpenTPU design into custom boards, providing PCB layouts, drivers, runtime libraries, and benchmark suites for rapid deployment.
  • Lowers the barrier for hobbyists and startups to experiment with an open TPU.

Details

Key Value
Target Audience Hobbyists, researchers, small companies wanting to build or test open TPU‑based accelerators
Core Feature Gerber PCB files, Linux kernel driver, C/C++ inference API, Dockerized benchmark for Llama/Qwen models
Tech Stack KiCad (PCB), Verilog/OpenTPU RTL, Rust/C for runtime, Docker, Python benchmark scripts
Difficulty Medium
Monetization Hobby (open source) or Revenue-ready: paid support/tiered consulting packages

Notes

  • fsbonetto repeatedly mentions OpenTPU and asks about hardware details (“Its a datacenter decommissioned board…”) – this kit gives them a ready‑to‑use path.
  • Enables the community to run the recursive self‑improvement loop on accessible hardware, as discussed in the thread.

Read Later