Our summary
Trying a puzzle idea can take much longer than inventing it. Marcin Dudek describes using coding agents to suggest ideas, then asking Jev to rank them before building anything in the browser. His first judge gave similar scores to almost everything and failed to help.
He replaced two broad questions with twelve checks that included clues, layout and earlier failed attempts. Code combined the answers with fixed weights. He then froze the questions and weights, keeping the agent that invented ideas from changing how those ideas would be judged.
Dudek reports that an idea ranked twelfth led to a useful hint and then a solution. Jev had not predicted it was the answer. This is one personal account, not a controlled comparison, and its combined scores should not be read as measured chances of solving a puzzle.
Key takeaways
- Check whether scores actually separate ideas before trusting whichever one ranks first.
- Include known clues and failed attempts, then combine separate judgments in inspectable code.
- Freeze the judge and keep idea generation separate so the generator cannot tune its own marks.