Our summary
A browser test may need to recognize a signed-in profile even when its greeting changes. Savan Vaghani pairs Playwright, which clicks and types, with Jev, which judges supplied text. Ordinary code still checks exact addresses, counts and rules; AI handles questions about meaning.
The application sends a small text outline of the page and asks whether it shows a profile or which listed control fits a goal. Jev cannot see screenshots or perform clicks. Record its suggestions without changing test results first, then compare them with known passing and failing cases.
The author reports matching answers on thirteen passing login checks, with one run per engine. Claude read much more context and ran through Claude Code, so the cost ratios are not an equal-input comparison. The study does not show how well either system detects broken pages.
Key takeaways
- Use code for exact addresses and counts; use a text judgment only where wording can vary without changing meaning.
- Include none or other among the answers, and measure mistakes on failing pages before adopting a pass mark.
- Save the supplied page text, questions, probability and model version; page content can mislead the model.