Our summary
An AI assistant may need to choose a helper or decide whether a proposed change fits the request. Vexjoy describes replacing some written AI judgments with Jev's set answers and probabilities. This makes the decisions easier for software to use, without proving that the judgments are correct.
The toolkit first uses code to search, count and enforce fixed rules. Jev then answers smaller questions about supplied evidence, such as whether a change adds unrequested behavior. In routing, one pass narrows the available helpers and another checks the smaller list. A writing model produces new text when the task needs it.
Vexjoy has not established that the probabilities deserve trust on these tasks. Reported routing comparisons used successive example sets, and shortening evidence can hide important details. The inspected code sends judgment material to hosted providers; secret-removal checks do not guarantee privacy. Model judgments should not replace permissions or firm safety rules.
Key takeaways
- Split a broad question into specific behaviors, and include examples of what should and should not count.
- Use code for exact checks and enforce actions separately from the model's judgment. A recommendation is not the repair itself.
- Check probabilities against labeled outcomes from your own work. Removing an untested cutoff helped this router, but does not show that every cutoff should disappear.