JevMade

Sign in
← Back to experiments

Apps & data pipelines

JEV_sees

JEV Sees adds a visual front end to JEV, converting local YOLO, CLIP, and RGB-D perception into scene text for structured decisions.

Bookmark: JEV_sees Keep this in your collection.
Leave a noteWhat would you try with this? : JEV_sees

Only you can see your notes.

Source screenshot of JEV_sees
SOURCE SCREENSHOTFull screenshot ↗

What it does

JEV receives scene text rather than pixels, relying on local perception models like Florence-2 and CLIP to build a text-based scene state. It is an experimental v0.1.0 release with no real-world safety or accuracy conclusions, and it enforces a strict 31,000-token hard limit on the generated scene state.

How you can use it

To use JEV Sees, you need a camera or video file and a JEV API key. The system tracks objects in the video and turns them into a text description that JEV can understand.

Start with the documented video example that asks a question about each visible pedestrian and returns a probability. Treat these as experimental scene judgments, not a safety detector: the local camera interpretation can be wrong, and Jev sees its text description rather than the original image. Do not rely on it for real-world pedestrian safety.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.