JEV receives scene text rather than pixels, relying on local perception models like Florence-2 and CLIP to build a text-based scene state. It is an experimental v0.1.0 release with no real-world safety or accuracy conclusions, and it enforces a strict 31,000-token hard limit on the generated scene state.
How you can use it
To use JEV Sees, you need a camera or video file and a JEV API key. The system tracks objects in the video and turns them into a text description that JEV can understand.
Start with the documented video example that asks a question about each visible pedestrian and returns a probability. Treat these as experimental scene judgments, not a safety detector: the local camera interpretation can be wrong, and Jev sees its text description rather than the original image. Do not rely on it for real-world pedestrian safety.