Project ideas from Hacker News discussions.

Show HN: Parseable, an open observability datalake, handles 100M time-series/min

📝 Discussion Summary (Click to expand)

Theme 1 – Scalability & Performance (handling massive cardinality and ingestion rates)
Commenters focused on how many time series the system can sustain, the ingest throughput, and how it compares to existing TSDBs.
- “scrape interval is 15s and sustained ingestion we have seen is ~3M samples/sec that is ~300 TB/day of raw ingest payload… The 100M figure is total unique series seen over time.” – nikhil4usinha
- “I thought Thanos and some other Prometheus variants can handle about 100 million active time series… Have you not pushed it past 100 million?” – goldeneye13_
- “We haven’t yet tried pushing it to the scale of billions yet. The max that we’ve gone to is 150‑180 million.” – parmesant

Theme 2 – Columnar Architecture (Arrow/Parquet, labels as columns)
The core technical differentiators were repeatedly highlighted: using Arrow for in‑memory processing, Parquet for durable storage, and treating labels as ordinary columns rather than per‑series indexes.
- “Our architecture is built around columnar design, and we use Apache Arrow for in‑memory columnar processing and Apache Parquet for durable columnar storage on S3‑compatible object storage.” – yashdotrv
- “labels are just columns in Parquet, so there is no per series index that grows with cardinality… What drives cost for us is ingestion rate (data points/s) and how much data a query has to scan for a particular time range not series count.” – nikhil4usinha

Theme 3 – Observability for AI Agents & Data Governance
Several participants noted the emerging need to trace LLM agents, tool calls, prompts, costs, etc., and stressed keeping that telemetry under the team’s own storage and compliance controls.
- “Also, one thing we’ve been thinking about a lot is how observability changes as agents become part of day‑to‑day engineering workflows… Observing them matters just as much as observing any other system. But it is equally important to decide where that telemetry data should reside. Our view is that teams should be able to keep these observability data close to them: in their own object storage, under their own retention, access, and compliance controls.” – yashdotrv
- “How does this any different then Iceberg?” – usernametaken29 (reflecting interest in storage‑layer choices for agent telemetry).


🚀 Project Ideas

Generating project ideas…

Observability Ingestion Capacity Planner

Summary

  • A web‑based calculator that estimates required ingestor count, storage size, and query cost for an observability data lake based on time‑series count, update rate, label cardinality, and retention period.
  • Core value proposition: gives platform engineers instant, data‑driven sizing insights to avoid over‑ or under‑provisioning when adopting columnar observability stores like Parseable.

Details

Key Value
Target Audience SREs, platform engineers, and architects evaluating observability stacks
Core Feature Interactive scenario modeling (adjust series, samples/sec, cardinality, retention) → outputs ingestor specs, storage GB/day, query scan cost
Tech Stack Frontend: React + TypeScript; Backend: Python/FastAPI; Computation: lightweight analytical model (can call Parseable’s stats API); Deployment: Docker/Kubernetes
Difficulty Medium
Monetization Revenue-ready: Subscription (tiered by monthly active users & API calls)

Notes

  • HN users asked for concrete numbers: “what's the data rate for each time series it can handle?” (msandford) and “sustained ingestion we have seen is ~3M samples/sec” (nikhil4usinha). This tool directly answers those questions.
  • Enables discussion on trade‑offs between ingestor scaling vs. query performance, a frequent topic in observability sizing threads.

Object Storage Retention & Compliance Manager for Observability Data

Summary

  • A policy‑driven service that automatically applies retention, encryption, and access‑control rules to Parquet observability data stored in users’ own S3‑compatible buckets.
  • Core value proposition: lets teams keep observability data close to them while meeting GDPR, SOC 2, and internal compliance without manual bucket‑level scripting.

Details

Key Value
Target Audience DevOps & security teams managing observability lakes in self‑hosted object storage
Core Feature Policy‑as‑code engine (YAML/JSON) that tags, expires, or encrypts Parquet partitions based on time, label values, and data sensitivity
Tech Stack Backend: Go or Rust (binary); Storage interface: AWS S3 SDK / MinIO; Policy engine: Open Policy Agent (OPA) or custom; Deployment: Helm chart for Kubernetes
Difficulty Medium‑High
Monetization Revenue-ready: Usage‑based pricing (per GB managed per month)

Notes

  • Commenters emphasized keeping data “in their own object storage, under their own retention, access, and compliance controls.” (Parseable’s founding team) and asked about cost and retention implications.
  • Provides a concrete way to implement those controls, sparking conversation on best practices for observability data governance.

High‑Cardinality Series Explorer

Summary

  • A lightweight web UI that lets engineers browse, filter, and sample high‑cardinality labels and series in a columnar observability lake (e.g., Parseable) using Arrow/Parquet column statistics for fast faceted search.
  • Core value proposition: turns opaque high‑cardinality metrics into an interactive, searchable view without scanning raw data, speeding debugging and exploration.

Details

Key Value
Target Audience Engineers debugging high‑cardinality metrics, SREs, and data analysts
Core Feature Faceted filter UI (label keys/values) + time‑range selector → returns sampled series and distribution histograms; leverages Parquet min/max/page stats for sub‑second responses
Tech Stack Frontend: React + Ant Design; WASM: Arrow.js + Parquet‑WASM reader; Backend: Go service serving pre‑computed stats & sample endpoints; Deployment: Docker
Difficulty High
Monetization Hobby (open‑source) – can be complemented with paid support/consulting

Notes

  • Users wondered about scaling to billions of series and how Parseable compares to TSDBs: “I would have expected your solution to scale to billions.” (goldeneye13_) and “Low level how does this compare to Victoriametrics…” (nwmcsween). This explorer gives immediate visibility into cardinality distribution, feeding those discussions.
  • Enables practical utility: quickly spot outliers, validate label cardinality assumptions, and share findings with teammates.

Read Later