Project ideas from Hacker News discussions.

OpenAI withdraws three mathematical results

📝 Discussion Summary (Click to expand)

Six prevalent themes in the HN discussion

  1. Low‑quality, rushed output – Many commenters described the AI‑generated papers as sloppy, hallucinated, or unreadable.

    “Early sentiments are a lot of the write ups still read like slop and it feels very rushed and not very polished.” — sashank_1509

  2. Incomplete or questionable formal verification – Only a fraction of the results have Lean formalizations, and even those may be semantically off or produced by fuzzing.

    “Only a subset contain Lean formalizations. And even for that subset, there's the potential that the formalization is semantically off (that is, it's a formalization for a slightly different problem).” — AlanYx

  3. PR/IPO‑driven release – The dump is seen as a publicity stunt aimed at boosting OpenAI’s image ahead of an IPO rather than a careful scientific contribution.

    “Is it a PR move designed for maximum IPO impact before actual mathematicians find errors and they have to withdraw many more…” — treebeard901

  4. Unpaid verification burden on the community – Critics argue OpenAI is shifting the work of checking proofs onto mathematicians without compensation.

    “They are sharing unproven work for PR, forcing the mathematicians community to do the verification job for them.” — mrpopo

  5. Limited practical impact – Even when correct, many results are viewed as theoretical curiosities (e.g., galactic algorithms) with little real‑world use.

    “It is an example of algorithm that is theoretically faster, but not with our sizes and hardware optimisations…” — kortzeus

  6. Threat to human expertise and jobs – Widespread concern that AI will replace mathematicians and engineers, eroding the pipeline of senior talent.

    “Without junior engineers how will there be senior software engineers in the future?” — dormento


🚀 Project Ideas

LeanProof Assistant

Summary

  • Automatically translates natural language math proofs into Lean 4 formalizations and runs them through the Lean checker to detect errors or incomplete steps.
  • Provides instant feedback to authors and reviewers, reducing the burden of manual verification and increasing trust in AI‑generated results.

Details

Key Value
Target Audience AI research labs, mathematicians, and proof‑assistant enthusiasts who want to validate LLM‑generated mathematics
Core Feature Natural‑language‑to‑Lean translation pipeline + automated Lean proof checking with error localization
Tech Stack Python, Lean 4, LLM API (e.g., OpenAI/GPT‑4o), Docker, FastAPI backend, React frontend
Difficulty Medium
Monetization Revenue‑ready: SaaS subscription per project or per‑proof‑check ($10/mo base, $0.01 per check)

Notes

  • HN commenters noted the lack of formalization: “Only a subset contain Lean formalizations… the formalization is semantically off” (AlanYx) and “If they had humility… they could have had humility in their announcement.” (SequoiaHope)
  • Enables community‑driven verification while giving labs a responsible way to share results before claiming breakthroughs.

MathPreprint Review Hub

Summary

  • A collaborative platform where researchers can post AI‑generated math preprints, annotate them, flag potential errors, and track verification status in real time.
  • Turns the ad‑hoc Twitter/arXiv discussion into a structured review process, giving credit to verifiers.

Details

Key Value
Target Audience Mathematicians, graduate students, and AI labs seeking community feedback on LLM‑produced mathematics
Core Feature Versioned preprint repository with inline commenting, error tagging, and reputation scores for reviewers
Tech Stack Node.js, Postgres, GraphQL, Git‑based storage, OAuth (GitHub/Google), React
Difficulty Medium
Monetization Hobby

Notes

  • Addresses the complaint that “they are sharing unproven work for PR, forcing the mathematicians community to do the verification job for them” (mrpopo).
  • Provides a venue for the kind of scrutiny that users like JohnKemeny wished for: “They are publishing proofs in natural language… and are not submitting to journals.”

AI‑Proof Auditor (Hallucination Detector)

Summary

  • An LLM‑based service that scans AI‑generated math texts for logical inconsistencies, unsupported claims, and hallucinated references, outputting a confidence score and highlighted dubious passages.
  • Helps readers quickly assess whether a paper is likely sound before investing time in deep reading.

Details

Key Value
Target Audience Researchers, peer reviewers, and journalists who need to triage large volumes of AI‑generated math content
Core Feature Hallucination detection model fine‑tuned on math corpora, citation verification, and consistency checking against known theorems
Tech Stack HuggingFace Transformers, FAISS vector store, SciBERT, Python Flask API, Streamlit demo
Difficulty High
Monetization Revenue‑ready: API usage‑based pricing ($0.005 per 1k tokens audited)

Notes

  • Responds to concerns like “It was an unreadable mess with some strong smells… it'd take a decent amount of labor to validate it” (IsTom) and “If you’re told everything you do is going to change the world… you’re ignorant of the output… then that’s a lot of hot air” (illwrks).
  • Gives a quick sanity check that could prevent the spread of “slop” before it consumes community effort.

Responsible Release Checklist Service

Summary

  • A lightweight web tool that AI labs can run before publishing math results, enforcing a checklist (internal Lean formalization, independent review, clear uncertainty statements, versioned artifacts).
  • Generates a verifiable attestation badge that can be displayed alongside the release.

Details

Key Value
Target Audience AI research teams, product managers, and compliance officers at companies like OpenAI
Core Feature Interactive checklist with automated checks (e.g., detects missing Lean files, prompts for uncertainty disclaimer) and PDF attestation generation
Tech Stack Static site (Next.js), serverless functions (Vercel), IPFS for artifact storage, optional blockchain attestation
Difficulty Low
Monetization Hobby

Notes

  • Directly tackles the frustration expressed by “They just blindly published results produced by the LLM, with 0 due diligence.” (autuni) and the call for “humility in their announcement.” (SequoiaHope)
  • Provides a tangible way for labs to demonstrate responsible practices, addressing the desire for transparency seen in the thread.

MathSlop Filter Browser Extension

Summary

  • Extension that scores incoming math‑heavy web pages (PDFs, arXiv entries, blog posts) for readability, formalization level, and language quality, overlaying a badge (e.g., “Low Slop”, “Medium Slop”, “High Slop”).
  • Lets users filter out low‑quality AI‑generated content before reading.

Details

Key Value
Target Audience Mathematicians, students, and anyone who frequently browses AI‑generated math discussions
Core Feature Heuristic scoring model (sentence length, jargon density, LaTeX vs plain text, presence of Lean/Coq snippets) + UI badge
Tech Stack JavaScript (WebExtension), Rust/Wasm for scoring, optional TensorFlow.js model, storage via chrome.storage
Difficulty Low
Monetization Hobby

Notes

  • Mirrors complaints such as “the write ups still read like slop and it feels very rushed and not very polished.” (sashank_1509) and “It was an unreadable mess with some strong smells.” (IsTom)
  • Gives readers an immediate cue to avoid wasting time on poorly written AI output, echoing the desire for better presentation.

Formalization Bounty Marketplace

Summary

  • A platform where AI labs post bounties (crypto or fiat) for formalizing specific AI‑generated theorems in Lean, Coq, or Isabelle; mathematicians claim bounties upon successful submission and verification.
  • Aligns incentives: labs get verified results, contributors earn rewards, and the community gains trusted formal proofs.

Details

Key Value
Target Audience Mathematicians interested in earning rewards for formal proof work, and AI labs needing verified results
Core Feature Bounty board, submission pipeline (upload Lean project), automated verification via CI, reputation and payout system
Tech Stack Solidity smart contracts (Polygon), IPFS for code storage, Lean 4 CI (GitHub Actions), React frontend
Difficulty High
Monetization Revenue‑ready: 5% platform fee on each bounty payout

Notes

  • Addresses the lean formalization gap highlighted by “Only ~42% of the posted results now have formalized proofs” (autuni) and the desire for community verification: “If they had humility… they could have had humility in their announcement.” (SequoiaHope)
  • Creates a market incentive that turns the verification burden into an opportunity, satisfying the call for “share the tool with mathematicians” (mrpopo).

Read Later