Project ideas from Hacker News discussions.

Jeeves. Reasoning improves Jev-like decision models

📝 Discussion Summary (Click to expand)

Theme 1 – “Jev” (decision‑model) is valued for being fast, cheap, and providing calibrated probabilities
- “I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean.” – doginasuit
- “…the beauty of Jev is that it is dirt cheap and insanely fast.” – sharih
- “Jev‑like models give calibrated decision probabilities…” – esafak

Theme 2 – Skepticism about the novelty or usefulness of the term “noul/decision model” and concerns about hallucinations
- “I think, it's a bit much to call this 'no hallucinations'. Technically true, but in practice you could still choose the wrong result or the probabilities can be off.” – k__
- “If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this.” – user3939382 (followed by discussion that it's not truly new)
- “In Bayesian statistics that’s called credence. Weird that they felt the need to invent a new term.” – LudwigNagasena

Theme 3 – Nostalgia and jokes about the revival of the “Ask Jeeves” brand
- “If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia.” – fishfasell
- “Ask Jeeves - Only took us 30 years to come full circle.” – thm
- “I met one of the founders once in Oakland. Amazing fella.” – victordmor (followed by many reminiscences)

These three themes dominate the conversation: praise for Jev’s speed/cost/probability output, doubt about the term’s originality and hallucination‑free claim, and playful references to the Ask Jeeves comeback.


🚀 Project Ideas

JevKit

Summary

  • A open‑source toolkit for fine‑tuning small language models into fast, calibrated decision models (noul) for binary classification tasks.
  • Enables developers to swap heavy LLMs with sub‑10 ms inference while preserving probability calibration.

Details

Key Value
Target Audience ML engineers and product teams building low‑latency classification services (spam detection, content moderation, feature flags)
Core Feature LoRA/QLoRA fine‑tuning pipeline that outputs a calibrated probability head, exports to ONNX/TorchScript, includes a prompt‑caching wrapper
Tech Stack Python, HuggingFace Transformers, PEFT, bitsandbytes, ONNX Runtime, FastAPI (optional serving)
Difficulty Medium
Monetization Revenue-ready: hosted API with usage‑based pricing (free tier up to 1M calls/mo)

Notes

  • HN commenters asked for good open source decision models trainable on own data (zerop) and wanted cheap/fast alternatives to LLMs (HarHarVeryFunny).
  • Provides a concrete way to satisfy those requests, sparking discussion on calibration techniques and latency trade‑offs.

JevBench

Summary

  • A benchmark suite and leaderboard that evaluates decision models (Jev‑style, LLMs with structured output, tiny classifiers) on accuracy, latency, and calibration metrics.
  • Gives teams an objective way to compare cost‑efficiency before integrating a decision model into production.

Details

Key Value
Target Audience Researchers, ML practitioners, and product managers evaluating decision‑model trade‑offs
Core Feature Automated runs of standardized tasks (e.g., irony detection, spam classification, Pokemon benchmark) with Brier score, ECE, p90 latency, and cost estimates
Tech Stack Python, pytest, Docker, HuggingFace Eval, Streamlit for leaderboard, GitHub Actions for CI
Difficulty Low
Monetization Hobby (open source, community‑driven)

Notes

  • Commenters discussed benchmarking Jev vs other models (TN1ck, Naitik88) and lamented lack of fair comparisons (mxkuzn, captainbland).
  • A public leaderboard would fuel HN debates and help teams pick the right model for low‑latency use cases.

JevVision

Summary

  • An API service that turns a multimodal prompt (image + text) into a calibrated binary decision probability, using a tiny vision‑language model fine‑tuned for decision tasks.
  • Delivers sub‑50 ms inference on GPU/CPU, letting developers replace heavy VLMs in real‑time moderation, safety checks, or UI feature flags.

Details

Key Value
Target Audience Developers building real‑time content‑moderation, spam‑filtering, or adaptive UI systems that need image‑aware decisions
Core Feature Pre‑built vision‑language decision model (e.g., Phi‑3‑Vision + LoRA) with probability head, prompt‑caching, and fallback CPU path
Tech Stack Python, Torch, HuggingFace Transformers, Triton Inference Server, FastAPI, optional TensorRT for acceleration
Difficulty High
Monetization Revenue-ready: pay‑per‑call ($0.0005 per 1k requests) with free tier

Notes

  • Swingboy asked for Jev‑like classifiers that support image input; HN users highlighted the need for fast multimodal decisions (zihotki, teravor).
  • Provides a ready‑to‑use solution that could spark discussion on multimodal calibration and edge deployment.

Read Later