Our summary
GC AI gave its team a day to try Jev, an AI model that judges supplied information instead of writing replies. Bardia Pourvakil describes 38 projects and experiments, from checking legal citations to sorting alerts and caring for a shared Slack pet.
The useful pattern was to ask small, concrete questions, then let ordinary code handle the action. Clean the text before asking, supply exact facts such as counts yourself, and test when uncertain answers need another model or a person. Several questions about one item differ from packing many similar records together.
These are hackathon reports, not results repeated by JevMade. A test-quality experiment improved its ranking score after replacing one broad question with specific checks, but the companion describes remaining failures and duplicate examples. Most reported tests used small, often synthetic sets and usually one run; check your own cases before relying on them.
Key takeaways
- Replace a broad judgment with concrete checks, and give choices a none or not-applicable option when needed.
- Let code calculate dates and counts, find extraction candidates and enforce permissions; a model judgment does not replace these rules.
- Tune cutoffs on labeled examples and check separate examples. Questions about one item can share a request; similar records may need separate calls.