Project ideas from Hacker News discussions.

RIP, vector database

📝 Discussion Summary (Click to expand)

Three dominant themes in the discussion

  1. Page load performance – Commenters reported wildly different experiences, ranging from instant loads to multi‑second delays, often attributing the variance to server load, connection throttling, or local ad‑blocking.
  2. OutOfHere: “It would be nice to have a page that actually loads. This one doesn’t. RIP.”
  3. alexjplant: “Takes 11 seconds to load on Firefox on Linux with 3G‑level throttling enabled in Dev Tools.”
  4. throwawy0352: “Loads really fast for me. (MacBook Air, average internet) … I got a Lighthouse score of 99 in Chrome.”
  5. wilj: “It has a pagespeed insights score of 55 and noticeably sluggish on my m3 max.”

  6. Vector‑DB vs. SQL‑based vector support – A sizable technical thread debated whether a dedicated vector database is necessary, with many arguing that adding native vector capabilities to traditional relational stores (SQL, Postgres, MySQL) avoids write amplification and simplifies operations.

  7. sreekanth850: “I find very little reason to use a pure vector database for enterprise retrieval. We built an enterprise retrieval engine on top of a SQL database with native vector support…”
  8. gopalv: “Your design choice went from a Postgres design pattern to a Mysql one… Postgres optimized for lookup and Mysql does for indexing on writes… This is a direct parallel to how Postgres and Mysql built indexes.”
  9. malisper: “Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums… the cost of the indirection is much smaller than it may first appear.”
  10. benesch: “An automatically generated internal ID: (segment ID, doc ID). The user‑provided primary key … turns into a secondary index at the storage layer.”

  11. Concerns about LLM‑generated or shill comments – Several users questioned the authenticity of recent posts, pointing to throwaway accounts, possible AI‑generated content, and the impact on community quality.

  12. phoghed: “fucking shills, making helpful comments and promoting seemingly nothing, what's this place coming to?”
  13. throwawy0352: “Yes, I get big money from the Internet Archive to promote their services. It's the new scheme that shills like me go for.”
  14. 0c3ca83: “Many of the commenters on this site are also obviously LLMs. I'd imagine that quite a few of the entities quoting aren't necessarily people…”
  15. FLeXMurphy: “Personally I haven't seen it too often; the other aspect is that HN is a forum for startups to pitch shit to each other, so this has been happening with or without marketers…”

These three threads—performance perception, the vector‑database architecture debate, and worries about AI‑driven or inauthentic commentary—dominated the conversation.


🚀 Project Ideas

EdgeSnapshotter

Summary

  • Generates static HTML snapshots of web pages during traffic spikes and serves them via a global edge CDN, reducing origin load and improving page load times.
  • Core value proposition: Faster, more reliable page delivery under variable traffic without manual scaling or caching configuration.

Details

Key Value
Target Audience Developers, bloggers, and SaaS owners experiencing slow load times or server overload during traffic bursts
Core Feature Auto‑detects high request rates, renders a snapshot with headless Chrome/Playwright, invalidates on content changes, and serves via CDN
Tech Stack Go/Rust worker, AWS Lambda or Cloudflare Workers, Playwright for rendering, Cloudflare/Fastly CDN, Redis for deduplication
Difficulty Medium
Monetization Revenue-ready: $0.01 per 1k snapshot views + bandwidth

Notes

  • Addresses OutOfHere’s wish: “It would be nice to have a page that actually loads” and alexjplant’s 11‑second load on throttled 3G.
  • Provides a practical way to mitigate server‑side bottlenecks highlighted by comments about traffic load mattering.

VecSQL

Summary

  • Postgres extension that stores vectors separately and uses an ANN index pointing to primary‑key IDs, enabling hybrid SQL + vector queries without the write‑amplification of traditional vector indexes.
  • Core value proposition: Simplifies retrieval pipelines by keeping vectors inside the relational store while eliminating costly index updates on vector changes.

Details

Key Value
Target Audience Engineers building RAG, semantic search, or any app needing combined relational and vector search
Core Feature ANN index (HNSW/Faiss) that stores only doc IDs; vector updates affect only the vector store, not the index, mirroring MySQL’s stable‑ID approach
Tech Stack C extension for Postgres, HNSWlib or Faiss bindings, SQL/MM interface, optional Docker image for deployment
Difficulty High
Monetization Hobby

Notes

  • Directly responds to sreekanth850’s preference for “SQL database with native vector support” and gopalv’s discussion on write amplification and index indirection.
  • Would appeal to HN commenters frustrated with managing separate vector DBs and syncing two systems.

LLMCommentDetector

Summary

  • Lightweight moderation tool that scores the likelihood a forum comment was generated by an LLM, flagging it for review via a browser extension or server‑side API.
  • Core value proposition: Helps communities maintain signal quality by automatically surfacing low‑effort, AI‑generated comments without invasive gated registration.

Details

Key Value
Target Audience Moderators, community managers, and platform developers (HN, Reddit, Lobste.rs, etc.)
Core Feature Extracts stylistic and linguistic features, runs a tiny ONNX/TensorFlow.js model to output an LLM‑likelihood score, integrates with existing moderation queues
Tech Stack Python/FastAPI backend, ONNX model, optional TensorFlow.js front‑end, Redis for caching, WebExtension for Chrome/Firefox
Difficulty Medium
Monetization Revenue-ready: $9 per month per moderator seat

Notes

  • Echoes 0c3ca83’s observation: “Many of the commenters on this site are also obviously LLMs” and the debate over gated registration.
  • Offers a less restrictive alternative to invite‑only systems while still curbing automated spam, a topic that generated strong discussion on HN.

Read Later