JevMade Sign in
← Back to guides

JevMade field notes / Written guide

Improve Jev questions with examples labelled by people

jev-align is an experimental command-line tool for improving a Jev decision function from examples. It looks for uncertain cases, asks a person to supply labels, and proposes changes to the function's written definition instead of silently treating model output as ground truth.

Original by SutroEvaluation

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

A decision function is saved instructions and answer options for one repeated judgment. jev-align selects uncertain examples for a person to label, then uses that evidence to suggest clearer wording instead of assuming the model’s first answers are correct.

The separate GEPA tool proposes revised wording from those labels. A user can add production examples and export the revised function. Any proposal should be tested on examples that were not used to create it before replacing the current version.

A revision can overfit—work well on its small training examples but poorly elsewhere. The project is experimental and live use calls paid outside services. People still define the task, protect sensitive examples, and decide whether separate test results show a real improvement.

Key takeaways

  1. Label uncertain examples instead of accepting them automatically.
  2. Evaluate revisions on examples not used to create them.
  3. Keep model keys and sensitive production data out of shared files.

This explanation follows the documented workflow.

GitHub project documentation

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.