JevMade

Sign in
← Back to experiments

Benchmarks & research

Jeeves (PostHog)

PostHog's Jeeves answers Jev-style decision questions, with an option to work through its reasoning before choosing.

Source screenshot of Jeeves (PostHog)
SOURCE SCREENSHOTFull screenshot ↗

What it does

The 9B model adapts Qwen3.5-9B, chooses among allowed answers and drafts reasoning in blocks of four tokens. It uses Jev's question format. PostHog reports better scores than Kev and Jev on its held-out and public JevBench tests, but weaker knowledge results on MMLU. Full reasoning takes much longer. Code is MIT-licensed and weights are Apache-2.0; Kev inspired the project. These are maker tests, not independently reproduced results.

How you can use it

Ask Jeeves which team should handle a customer's message, whether it needs urgent attention and how frustrated the customer is. Its estimates let your program choose among those answers. The README shows running it locally and asking for reasoning too; shorter reasoning limits reduce waiting. Substantial computing memory is needed.

Keep this for later

Sign in to bookmark experiments, guides and videos, and keep notes only you can see.

Continue to sign in

We’ll bring you back to this listing.