JevMade Sign in
← Back to guides

JevMade field notes / Hackathon lessons and public pattern matrix

Ask smaller questions and let your code act

Bardia Pourvakil and the GC AI team explain how narrow questions, clean source text and separate safety rules helped their Jev prototypes, while broad judgments and arithmetic exposed limits.

Original by Bardia Pourvakil and the GC AI teamAgent workflows

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

GC AI gave its team a day to try Jev, an AI model that judges supplied information instead of writing replies. Bardia Pourvakil describes 38 projects and experiments, from checking legal citations to sorting alerts and caring for a shared Slack pet.

The useful pattern was to ask small, concrete questions, then let ordinary code handle the action. Clean the text before asking, supply exact facts such as counts yourself, and test when uncertain answers need another model or a person. Several questions about one item differ from packing many similar records together.

These are hackathon reports, not results repeated by JevMade. A test-quality experiment improved its ranking score after replacing one broad question with specific checks, but the companion describes remaining failures and duplicate examples. Most reported tests used small, often synthetic sets and usually one run; check your own cases before relying on them.

Key takeaways

  1. Replace a broad judgment with concrete checks, and give choices a none or not-applicable option when needed.
  2. Let code calculate dates and counts, find extraction candidates and enforce permissions; a model judgment does not replace these rules.
  3. Tune cutoffs on labeled examples and check separate examples. Questions about one item can share a request; similar records may need separate calls.

The test-quality result needs caution. AUC measures how well a score ranks examples, not the share of correct answers: 0.72 does not mean 72% accuracy. The article says eight checks; its companion matrix says nine plus a count calculated by code. The matrix calls the project failed and reports 298 duplicate tests appearing on both sides of a test split. We have not resolved those differences or repeated the tests. GC AI reports 38 projects, with 23 submitted across three tracks. Most results used small, often synthetic sets and usually one run. Removing cookie dialogs and narrowing a liability question fixed separate causes of false alarms. Matrix mockups are illustrations, not verified screenshots; private project links and recordings were not opened. The article was published October 6, 2026; the matrix dates the hack day September 29.

GC AI · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.