Jev in production › Evaluation and testing

pytest-jev

A pytest fixture whose jev.expect sends every holds or lacks claim about an LLM app's reply to Jev in one request and fails the test on the returned probabilities. On 12 example tests it matched Claude Sonnet 5's verdicts and ran 110× cheaper (author).

Open on GitHub ↗

110× cheapermeasured against a baseline, as published by the source
Use
Evaluation and testing
Industry
Developer tools
Form
Open-source tool
Stage
In production
Plugs into
pytest
Listed
2026-09-22
Found via
discord
Repository
allebee/pytest-jev
Stars
7
Forks
1
Last push
2026-09-21
Language
Python
License
MIT
pytest-jev screenshot
docs/demo/demo.gif in the allebee/pytest-jev README, MIT; shown from GitHub.

The README opens with

Semantic assertions for pytest, judged by TypeSafe's Jev, a model that returns calibrated probabilities instead of text.

Same verdicts as Claude Sonnet 5 on the example tests, 5× faster and 110× cheaper. Benchmark · real runs of jev-1.13 through OpenRouter

All the claims about one text go to Jev in a single request. When a prompt change breaks the reply, the failure shows which claims broke and how sure Jev was:

It needs Python 3.10+ and pytest 7.4+. Without a key, tests that use jev are skipped (see CI). Every output in this README is from a real run of examples/ against jev-1.13.

Badge

For the project's own README, linking back here:

Listed in Jev in production

Also used for evaluation and testing