Project ideas from Hacker News discussions.

Rampart: Browser native on-device PII radaction

📝 Discussion Summary (Click to expand)

1. Distrust of government handling of personal data
- “The skeptic in me really can't trust our current government to protect my information.” – sgnelson
- “You could never trust the US government to protect our information since it moved online.” – Cider9986
- “So that’s why big ballz (co author of the OP) put our PII in an insecure AWS instance via DOGE’s starlink terminal.” – iAMkenough

2. Concerns about the effectiveness and technical limits of PII redaction solutions
- “98.4% is nowhere near good enough to call it PII redaction.” – nhinck2
- “The only thing that makes these people happy is zero data retention and deterministic redaction at the source. Regex over arbitrary string literals does not represent determinism in this context.” – bob1029
- “Is this good enough for HIPAA?” – Onavo
- “The lowest hanging fruit … is to clearly tell people … not to share any personal info with chatbots … the government should have extended HIPPA … to AI companies.” – dwa3592

3. Emphasis on transparency, open‑source availability, and the ability to run/inspect the model locally
- “Source available here … you can also switch that out for your own with the source.” – BowBun
- “I mean, it can be run locally, so you don't necessarily have to trust it, unless there's been any model‑as‑a‑vector CVE.” – goodmythical
- “The source code should be public domain, no?” – Cider9986


🚀 Project Ideas

Deterministic PII Redaction Middleware for LLMs

Summary

  • A plug‑and‑play library that intercepts prompts and responses to LLMs and applies deterministic, rule‑based PII redaction (SSN, DOB, address, etc.) before the model sees the data and after it generates output.
  • Core value: guarantees zero storage of raw PII and enables developers to offer AI‑powered chat without risking leaks, addressing the trust concerns raised by HN commenters.

Details

Key Value
Target Audience Developers building AI chatbots, SaaS platforms handling user‑submitted text, government contractors
Core Feature Deterministic regex‑plus‑contextual redaction engine that never retains input; optional local model switch for custom rules
Tech Stack Rust (core engine), Python bindings, WASM for edge deployment, optional integration with HuggingFace Transformers
Difficulty Medium
Monetization Revenue-ready: SaaS tiered pricing (free dev tier, paid per‑million‑tokens processed)

Notes

  • HN users emphasized “zero data retention and deterministic redaction at the source” (bob1029) and the need for “zero data retention” to trust government AI (sgnelson, dw3592). This middleware directly satisfies those asks.
  • Enables discussion on balancing utility vs. privacy; can be showcased in a demo redacting the example SSN shared in the thread while preserving conversation flow.

Zero‑Data‑Retention AI Inference Service (ZDRaaS)

Summary

  • A stateless API that runs LLMs inside secure enclaves (e.g., AWS Nitro Enclaves or confidential VMs) guaranteeing that neither prompts nor generated text are persisted to disk or logs.
  • Core value: offers the “magical AI experience” investors want while providing cryptographic proof of no data retention, mitigating the government trust gap highlighted in the discussion.

Details

Key Value
Target Audience Enterprises handling sensitive citizen data, health‑tech apps, fintech platforms needing HIPAA‑grade assurances
Core Feature Enclave‑based inference with automatic memory wiping after each request; audit‑log‑only metadata (request ID, timestamps)
Tech Stack Go/ Rust for enclave runtime, TensorFlow‑Serve or vLLM for model serving, attestation via Intel SGX/ AMD SEV, API gateway (Envoy)
Difficulty High
Monetization Revenue-ready: Pay‑per‑request + enclave usage (e.g., $0.0005 per 1k tokens)

Notes

  • Commenters like “sgnelson” and “dw3592” voiced distrust of government handling PII; a verifiable zero‑retention service would give them confidence to adopt AI for public‑service use cases.
  • Could spark HN debate on trade‑offs between performance and privacy, and attract interest from agencies looking to pilot compliant AI.

Video PII Redaction Studio

Summary

  • Desktop / web application that processes video streams in real time, automatically detecting and redacting faces, license plates, on‑screen text (e.g., SSNs displayed in forms), and audio PII using OCR and speech‑to‑text filters.
  • Core value: gives content creators, broadcasters, and government agencies an easy way to publish video without exposing personal information, directly answering the request for “a PII redaction model for video” (swiftcoder).

Details

Key Value
Target Audience Media companies, law‑enforcement agencies, corporate compliance teams, YouTubers publishing dashcam or meeting footage
Core Feature Real‑time pipelines: face detection (MTCNN/BlazeFace), license‑plate detection (YOLOv8), OCR (Tesseract + custom SSN regex), audio PII bleeping (Whisper + keyword spotting)
Tech Stack Python (OpenCV, PyTorch), Electron for desktop UI, WebAssembly + WebGPU for browser version, Docker for deployment
Difficulty Medium
Monetization Hobby (open‑source core) with optional paid Pro pack for GPU‑accelerated cloud processing and SLA support

Notes

  • Swiftcoder explicitly asked for a video PII redaction model; this tool satisfies that need and can be demonstrated on the example video of the DOGE terminal leak.
  • Provides practical utility for HN community members who often share screen recordings or dashcam clips and worry about inadvertent PII exposure, encouraging discussion on open‑source vs. commercial offerings.

Read Later