Three prevalent themes in the discussion
- Prompt quality is a major obstacle
- “The prompts provided are atrocious. It's amazing that the LLMs actually built something useful.” – smokel
- “It's so confusing how the actual article is in English, but the prompts are just gibberish.” – InsideOutSanta
- “Yeah, I'm thinking the nigh‑unreadable AI-speak we get these days makes a lot more sense if this is what they're training on. Or maybe the author has translated the prompts from another language?” – daemonologist
-
“Or having another model proofread the prompts and write clearer instructions.” – cavemandaveman
-
Models benefit from starting fresh each iteration rather than building on flawed outputs
- “I feel like the LLMs would have done better if they had started from scratch each time, rather than being burdened by the output from the previous attempt.” – InsideOutSanta
- “Testing the next model by iterating on the first model’s crappy foundation. Why not start from the original each time?” – aksss
-
“Instead of fixing a broken version, maybe each model should have started from scratch.” – tehlike
-
Observations of LLM performance versus human‑crafted solutions (and hardware feasibility)
- “Astra … had a highscore of around 1.7M … it disassembled parts of the ROM to extract information about the game.” – criemen
- “Astra … reacts to the visuals on‑screen, planning ahead by estimating velocity of all objects on screen… behaves more like a regular 'perfect' player.” – criemen (comparing Astra to the winning human solution)
- “You could actually play Prince of Persia on an 8088 with a CGA card!” – hapless (highlighting that the original game runs on very low‑end hardware)
- “Prince of Persia is a piece of contemplative, subtle, beautifully, artistically minimal motion puzzle art.” – dofm (appreciation of the game’s design amid the technical discussion)