Jev in production › Evaluation and testing
A pytest fixture whose jev.expect sends every holds or lacks claim about an LLM app's reply to Jev in one request and fails the test on the returned probabilities. On 12 example tests it matched Claude Sonnet 5's verdicts and ran 110× cheaper (author).

Semantic assertions for pytest, judged by TypeSafe's Jev, a model that returns calibrated probabilities instead of text.
Same verdicts as Claude Sonnet 5 on the example tests, 5× faster and 110× cheaper. Benchmark · real runs of jev-1.13 through OpenRouter
All the claims about one text go to Jev in a single request. When a prompt change breaks the reply, the failure shows which claims broke and how sure Jev was:
It needs Python 3.10+ and pytest 7.4+. Without a key, tests that use jev are skipped (see CI). Every output in this README is from a real run of examples/ against jev-1.13.
For the project's own README, linking back here: