Our summary
A program sometimes needs to choose an answer rather than write a reply. Yury Selivanov introduces a Python interface for asking Jev those questions. His two experiments explore a useful boundary: recognizing what someone is typing is not the same job as writing working code.
The first experiment asks whether unfinished input is Python code or English. The second asks Jev to choose pieces of a program's structure, while ordinary Python code supplies punctuation and spacing. A separate language model expands the instructions beforehand and reviews the finished code by inspection afterward.
Neither experiment worked as the author hoped. Partly typed code could look like English, and programs with valid spelling and structure were still mostly wrong. The generated programs were not executed; a model's review is not a test. This experimental interface and these examples need checking against the version in use.
Key takeaways
- Give Jev a bounded question: choose an option, rate something on a described scale, or estimate a yes-or-no probability.
- Incomplete input can be ambiguous. Test the actual typing patterns your interface will receive.
- Code that has the right structure can still do the wrong thing. A language model's inspection does not replace running appropriate tests.