Project ideas from Hacker News discussions.

Greg Kroah-Hartman – Security in the LLM Age [video]

📝 Discussion Summary (Click to expand)

Theme 1 – Overstated / misleading claims about LLM‑found bugs
- “Wild hype machine around these companies and uncritical parroting of every press release…” – usernomdeguerre
- “Widely proclaiming that your new model is so dangerous it needs to be released only to select people… widely proclaiming the model easily found 79 bugs… GKH says it took 1 hour to fix all of them because most weren't bugs…” – djoldman

Theme 2 – Human expertise remains essential; LLM output needs expert review
- “Reading through a Claude generated false positive is absolutely excruciating… takes hours of expert human labour to understand, test, and discard.” – OtherShrezzing
- “Leave them to their own devices at your peril. Trust nothing they do. Yet directly guide them, monitor everything they do, some value emerges.” – b112
- “Humans and especially human reviewers remain critical for the sustainability of software systems.” – asaiacai

Theme 3 – LLMs behave like inexperienced interns lacking real‑world judgment
- “Right now, all top tier LLMs are as eager, bright 20ish year old interns. Very gung ho, full of energy, loads of book learning, no real world experience…” – b112
- “Like a top percentile fresh grad on meth. Still a lack of real world experience plus some bizarre failures…” – fc417fc802


🚀 Project Ideas

Generating project ideas…

VulnTriage Assistant

Summary

  • An AI‑assisted triage platform that ingests LLM‑generated vulnerability reports, cross‑checks them against known CVEs, patch history, and code context to automatically filter false positives and rank genuine issues.
  • Core value: saves maintainers hours of manual review by delivering concise, actionable summaries with confidence scores.

Details

Key Value
Target Audience Open‑source project maintainers, security volunteers, dev‑sec teams
Core Feature Automated validation pipeline: static analysis + fine‑tuned LLM + CVE lookup → severity score & short explanation
Tech Stack Python, FastAPI, HuggingFace Transformers, PostgreSQL, Docker, optional WASM sandbox
Difficulty Medium
Monetization Revenue-ready: SaaS subscription (free tier up to 100 reports/mo, paid tiers $20‑$200/mo)

Notes

  • HN users complained about “reading through a Claude generated false positive is absolutely excruciating” and that “the majority of stuff is overly‑verbose nonsense which takes hours of expert human labour to understand, test, and discard.”
  • Provides a concrete way to turn the noisy LLM output into useful signal, sparking discussion on balancing AI assistance with human expertise.

KernelSpec Model Hub

Summary

  • A hosted service offering specialized LLMs fine‑tuned on the Linux kernel source, coding conventions, historical patches, and threat models, accessible via API for accurate vulnerability discovery and code review.
  • Core value: delivers higher‑precision bug findings by grounding the model in domain‑specific knowledge, reducing hallucinated reports.

Details

Key Value
Target Audience Security researchers, kernel developers, companies auditing low‑level code
Core Feature Fine‑tuned model + Retrieval‑Augmented Generation (RAG) over kernel codebase and CVE database
Tech Stack PyTorch, FAISS, HuggingFace Transformers, Kubernetes, GPU nodes, REST/GraphQL API
Difficulty High
Monetization Revenue-ready: pay‑per‑API‑call ($0.001 per 1k tokens) or enterprise license ($5k/yr)

Notes

  • Commenters noted that “a model trained on triage data to validate the first one's findings would be good” and that “specialized models trained on say Linux kernel specifics … can be much more relentless than humans.”
  • Enables the community to move from hype‑driven claims to measurable, reproducible results, inviting debate on model licensing and open‑source access.

SlopFilter CLI/GitHub Action

Summary

  • A lightweight command‑line tool (also usable as a GitHub Action) that runs LLM vulnerability reports through a rule‑based filter and a tiny ML classifier to auto‑discard obvious false positives such as “something crashed”, duplicates, or already‑fixed issues.
  • Core value: instantly cuts down noise in security scanning pipelines with virtually zero setup.

Details

Key Value
Target Audience Maintainers of open‑source repos, security hobbyists, CI/CD pipelines
Core Feature Pattern‑matching + lightweight ONNX classifier trained on known false‑positive patterns (e.g., from Mythos reports)
Tech Stack Rust (or Go), regex libraries, ONNX Runtime, optional GitHub Action wrapper
Difficulty Low
Monetization Hobby (open‑source MIT license) – optional hosted version with free tier and $5/mo for private repos

Notes

  • Users highlighted that “Most of the real low hanging fruit was picked up by humans years ago. When doing automated scanning, the majority of stuff is overly‑verbose nonsense…”, indicating a strong appetite for a filter that removes the “slop”.
  • Easy to adopt in existing workflows, likely to generate discussion on optimal rule sets and community‑maintained false‑positive databases.

Read Later