Our summary
A decision function is saved instructions and answer options for one repeated judgment. jev-align selects uncertain examples for a person to label, then uses that evidence to suggest clearer wording instead of assuming the model’s first answers are correct.
The separate GEPA tool proposes revised wording from those labels. A user can add production examples and export the revised function. Any proposal should be tested on examples that were not used to create it before replacing the current version.
A revision can overfit—work well on its small training examples but poorly elsewhere. The project is experimental and live use calls paid outside services. People still define the task, protect sensitive examples, and decide whether separate test results show a real improvement.
Key takeaways
- Label uncertain examples instead of accepting them automatically.
- Evaluate revisions on examples not used to create them.
- Keep model keys and sensitive production data out of shared files.