Theme 1 – Calibration and reliability of Jev’s probability outputs
- “On actual NLP problems (including ones with uncertainty under human review) it does appear to be well calibrated.” – cannedbread
- “Jev should ideally respond in a non‑random way, it should just list out the probabilities.” – jezzamon
- “Claims like this needs to be deeply analyzed… it’s hard to exactly say where there’s some internal mechanism generating a true answer and where we’re just getting lucky with some distribution.” – dvt
- “If I feed that into a model, the answer I want is: ‘the combination of the model and the provided state has nothing useful to add to your prior’.” – amluto
Theme 2 – Model’s inherent capabilities and limits (size, emergent learning, vs luck)
- “It’s a very small and dumb model. That’s the only reason why it’s so fast.” – singularity2001
- “Models do have some emergent capabilities… semantics is actually learned … some math seems like it also might be learned.” – dvt
- “Unless the numbers are all made up, this didn’t read like slop to me… maybe it’s your slopatron that’s miscalibrated!” – exe34
Theme 3 – Human analogy and expectations for model behavior (priors, uncertainty, pedantic answers)
- “Wouldn't asking humans have this same kind of problem too? … we also have our own biases.” – clarle
- “The author is expecting Jev to output the equivalent of 'H: 50%, T:50%', not the equivalent of 'TTHTTHTHH...'” – unholiness
- “The point of the theoretical problems is that they should be the easiest cases to handle. How can you trust the probabilities from real world classifiers if it can't even handle well defined problems.” – charcircuit
- “What I want out of a system like Jev is to tell me how the probabilities change as a result of the per‑sample data I provide.” – amluto