Three dominant themes in the discussion
| Theme | Supporting quote(s) |
|---|---|
| 1️⃣ Agentic tools are crossing the “can it be done” rubicon, but real‑world reliability is still shaky | > In the past year, agent harnesses crossed the “can it be done” rubicon. – hirvi74 “I triggered it once by accident … it broke everything. Now I just use ask mode, and even that is wrong half the time.” – al_borland |
| 2️⃣ Success hinges on explicit specs and human‑written tests; LLMs need clear guidance | “I saw a post … his spec document for the AI was 107 pages long. So maybe what I’m doing wrong is not giving the AI a literal novel of spec.” – al_borland > “Basically all the examples of LLM's building impressive things have been because they have human written tests to base the implementation on. If you have an LLM write the tests the results are far less impressive or valuable.” – slopinthebag |
| 3️⃣ Debate over LLM “reasoning” and safety (prompt‑injection, alignment) | “It helps to know that LLMs don’t ‘reason’… prediction is the training objective.” – theteapot > “LLMs are foundationally incapable of always and consistently preventing prompt injection attacks …” – hbcdbff (citing Anthropic data) |
These three themes capture the community’s focus on agentic maturity, the necessity of precise human‑crafted guidance, and the ongoing scrutiny of LLM reasoning and safety.