JevMade Sign in
← Back to guides

JevMade field notes / Commercial classification case study

Check whether a startup description names customers, a problem and income

Vincent Forat tests Jev on short invented startup descriptions in eight languages, using three yes-or-no questions and a keyword backup to flag information the description may be missing.

Original by Vincent ForatClassification

Listen to this guide

JevMade’s plain-English explanation

0:00 /

Our summary

A startup description can sound plausible without naming who needs it, what problem it addresses or how it earns money. Vincent Forat, founder of Preuve AI, tests whether Jev can flag these gaps. The article examines what the words say, not whether customers will actually buy the proposed product.

Forat asks three yes-or-no questions about each description. Probabilities of at least 0.8 count as yes, and at most 0.2 as no; a keyword check handles the middle range. He reports 326 correct answers among 332 clear decisions, and 349 correct decisions out of 384 for the combined method.

The descriptions are invented, short and clean, and the questions were tuned on the same set. One scenario's expected answer was later changed across all eight languages. The full data is not published, and the feature was not online when the article was written. This commercial case study does not establish accuracy with real users or a reliable ranking of languages.

Key takeaways

  1. Decide whether a description must state a problem directly or may imply it. Different interpretations change what counts as a correct answer.
  2. Count the backup's mistakes too. The roughly 98% result excludes 52 uncertain decisions; the full combined method reached 349 out of 384, or 90.9%.
  3. Test new, untuned descriptions from actual users before relying on the feature. A description naming customers and revenue is not evidence of demand.

The review and tables were read; the tests were not repeated. Forat reports a September 19 run on Jev version one point thirteen point zero: 16 scenarios in eight languages, 128 descriptions and 384 yes-or-no decisions. These return a yes probability, not a separate confidence value. Keywords handled 52 middle-range decisions; six clear decisions were wrong. After one expected answer changed across eight translations, keywords wrongly warned that information was missing 95 times, versus 27 for the combined method. They missed actual gaps 37 and eight times, respectively. The combined method reportedly got 72 out of 72 decisions right on another author-written set, but complete data and original results are not linked. Forat sells startup validation, discloses no TypeSafe affiliation and said the feature was not yet online. His small-run timings and costs were not independently verified.

Preuve AI · Original published

Read the original guide Opens the author’s site in a new tab.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.