Project ideas from Hacker News discussions.

Show HN: AI search for every photo and every frame of video on macOS

📝 Discussion Summary (Click to expand)

Theme 1 – Performance & optimization concerns
Users stressed that processing large video/photo collections hinges on sampling strategies and proxy workflows.
- “Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.” – hn3ufz62f7
- “Maybe you need a minimal downscale version as well, I heard is very common technique in the video editing world.” – pezgordo
- “Proxies. You transcode proxies from the original media, edit off those, then you use OM for the final render.” – Forgeties79

Theme 2 – Tool comparisons & ecosystem alternatives
Commenters compared the project to existing solutions (Immich, Apple Photos, Vision Framework) and suggested other models or native approaches.
- “possibly off-topic, but for anyone interested in this on a more cross‑platform / holistic basis, Immich does this … an approximate AI search for photos & videos” – lucideer
- “Tried immich and there was a lot not to like. It behooves everyone to try each one and see if it fits.” – mannyv
- “Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.” – postalcoder
- “Why chose CLIP to do this. Have you tried small VLMs like Qwen‑VL? I believe those models have video encoders can better perform at this scenario.” – yt1998

Theme 3 – Copyright, sherlocking & LLM implications
A sub‑thread debated whether LLMs could recreate functionality without infringing copyright, touching on the “sherlocking” concept.
- “Can you copyright things like this now that LLMs exist? … But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X … and then get round copying laws?” – alt227
- “Copyrights protect the literal text of your program and binaries. Not the design and functionality. Getting Sherlocked is having your functionality duplicated, independent of the code itself.” – robotresearcher
- “copywriting is the art of writing copy for products/marketing … maybe you were thinking of sherlocking” – tough


🚀 Project Ideas

SceneAware Video Indexer

Summary

  • Automatically detects scene changes in videos and samples representative frames at adaptive rates, then indexes them with CLIP embeddings for fast natural‑language search.
  • Generates low‑resolution proxies and caches thumbnails so browsing large libraries (e.g., 12k+ videos) feels instantaneous.

Details

Key Value
Target Audience Video editors, media librarians, cinematographers working with large video collections
Core Feature Scene‑based adaptive frame sampling + CLIP vector index + proxy generation
Tech Stack Python, FFmpeg, OpenCV, PyTorch/CLIP, FAISS, Electron/React (UI), Docker for deployment
Difficulty Medium
Monetization Revenue-ready: SaaS subscription ($10/user/mo) for hosted index, team sharing, and GPU‑accelerated re‑indexing

Notes

  • Users complained about “frame sampling rate is the whole ballgame” and needing “scene detection” to cut processing time (hn3ufz62f7, xnx).
  • A tool that “creates proxies… edit off those, then you use OM for the final render” would streamline workflows (Forgeties79).
  • Enables near‑instant searches like “find photos of kitchens” or “houses with palm trees” without waiting days for full‑frame processing (measure2xcut1x).

Network‑Aware Media Asset Manager (NAMAM)

Summary

  • Self‑hosted DAM that works directly over SMB/NFS shares, eliminating the need to ingest files into a proprietary library.
  • Provides natural‑language search via multimodal embeddings, on‑the‑fly proxy/thumbnails, and granular access controls.

Details

Key Value
Target Audience Creative teams, photographers, and studios storing assets on NAS or shared drives
Core Feature Transparent SMB access + CLIP‑based search + proxy/thumbnails generation
Tech Stack Node.js (backend), PostgreSQL + pgvector, Rust for FFmpeg‑based proxy workers, React + Ant Design (web UI), optional Electron desktop client
Difficulty Medium
Monetization Revenue-ready: Hosted offering with tiered storage ($15/user/mo) + optional enterprise support

Notes

  • Commenters noted the pain of “photos are on an external SMB network drive” limiting Apple Photos search (measure2xcut1x).
  • Desire for a “cross‑platform / holistic basis” solution like Immich but with better thumbnail handling and no forced folder duplication (lucideer, manyv).
  • Ability to “search for People / faces / pets” and “stock photography” queries aligns with the natural‑language search use case (qprofyeh, measure2xcut1x).

Embedded CLIP Sidecar Tool (ECSidecar)

Summary

  • CLI utility that extracts frames at adaptive intervals, computes CLIP embeddings, and writes them as XMP sidecar files (or embedded metadata) alongside original media.
  • Enables portable, offline‑first search: any machine can read the sidecars and perform similarity queries without a central database.

Details

Key Value
Target Audience Photographers, archivists, and developers who want vendor‑agnostic, portable AI tags
Core Feature Adaptive frame sampling → CLIP embedding → XMP/JSON sidecar generation
Tech Stack Rust (core), ffmpeg‑rs for frame extraction, torch‑rs or ONNX Runtime for CLIP, xmp crate for metadata handling
Difficulty Low
Monetization Hobby

Notes

  • Users expressed interest in “picture embeddings attached to the file by the camera” as a future ideal (freecodeio); ECSidecar realizes that today via sidecars.
  • The idea of “standard model like CLIP” being portable and not locking into a single vendor’s embedding was highlighted (collingreen).
  • Enables workflows where users can “just get gemini to write a DAM for me in a weekend” by providing the underlying portable metadata (lucideer).

Read Later